Limitation or failure · Biology · Chemistry · Materials science
LLM experiment planners fail to beat a statistical baseline across biochemical search tasks
A seven-dataset evaluation of frontier language models making sequential experimental choices under fixed budgets in biochemical optimization problems.
Summary
Researchers benchmarked five frontier language models on seven combinatorial datasets spanning protein engineering, reaction optimization, molecular design, peptide self-assembly and catalysis. Chemical context helped on average but with high variance, and no tested configuration decisively beat a mean statistical baseline across domains. Belief and action diagnostics showed that models intended to explore but behaved exploitatively when recent history remained in context.
AI role
Five frontier models selected experiments in Bayesian-optimization loops using combinations of literature priors, chemical context and accumulating observations.
Narrative role
Qualifies autonomous-lab claims with a cross-domain failure mode: verbal scientific intent and access to prior knowledge did not reliably translate into effective exploration or stronger experimental choices.
Caveat
This is a preprint evaluation over historical datasets and simulated sequential choices, not a prospective physical laboratory campaign. Results may depend on prompting, memory design and the selected statistical baseline.