Capability benchmark · Chemistry · Medicine
SMDD-Bench tests agents on multi-turn small-molecule drug design
A benchmark of guaranteed-solvable, long-horizon small-molecule drug-design tasks across protein targets and chemical task types.
Summary
SMDD-Bench introduced a standardized agent benchmark for real-world small-molecule drug design. It covers 502 multi-turn task instances across five design task types and 102 protein targets, with frontier models still solving less than half of the benchmark.
AI role
LLM agents were evaluated on planning, tool use, chemical and biological reasoning, and iterative design under limited oracle calls.
Narrative role
This adds a chemistry and medicine capability measure that is closer to computational drug-discovery work than static science QA, while also showing substantial remaining headroom.
Caveat
The benchmark is a preprint and evaluates computational task completion rather than wet-lab discovery or clinical utility.