Limitation or failure · Chemistry · Medicine · Computer science
R-GroundBench reveals a recognition-to-editing gap in patent molecular structures
A patent-derived benchmark measures R-group identification and generation of correctly edited molecular structures under different input modalities.
Summary
R-GroundBench tests molecular families represented by variable substituents, common in pharmaceutical patents. Strong multiple-choice recognition does not carry over to exact editing. In the full results table, GPT-5.5 reaches 66.1% on Basic Hard VQA with image input, but only 17.0% exact match on Hard image-to-SMILES generation; its Hard text-to-SMILES score is 38.4%. Chemical-domain models also show substantial failures.
AI role
Language and vision-language models grounded variable substituents and generated molecular SMILES from Markush structures and instructions.
Narrative role
Provides a concrete chemical-grounding failure behind apparently strong recognition scores, relevant to agents interpreting patents and constructing molecular candidates.
Caveat
Preprint with constructed instructions and exact-match grading, not prospective drug discovery. Abstract statements that visual generation stays below 8% conflict with the table’s best 17.0%; this entry uses the table and keeps task and modality boundaries explicit.