Limitation or failure · Chemistry · Medicine · Computer science

R-GroundBench reveals a recognition-to-editing gap in patent molecular structures

A patent-derived benchmark measures R-group identification and generation of correctly edited molecular structures under different input modalities.

Summary

R-GroundBench tests molecular families represented by variable substituents, common in pharmaceutical patents. Strong multiple-choice recognition does not carry over to exact editing. In the full results table, GPT-5.5 reaches 66.1% on Basic Hard VQA with image input, but only 17.0% exact match on Hard image-to-SMILES generation; its Hard text-to-SMILES score is 38.4%. Chemical-domain models also show substantial failures.

AI role

Language and vision-language models grounded variable substituents and generated molecular SMILES from Markush structures and instructions.

Narrative role

Provides a concrete chemical-grounding failure behind apparently strong recognition scores, relevant to agents interpreting patents and constructing molecular candidates.

Caveat

Preprint with constructed instructions and exact-match grading, not prospective drug discovery. Abstract statements that visual generation stays below 8% conflict with the table’s best 17.0%; this entry uses the table and keeps task and modality boundaries explicit.