Limitation or failure · Biology · Chemistry · Materials science

LLM experiment planners fail to beat a statistical baseline across biochemical search tasks

A seven-dataset evaluation of frontier language models making sequential experimental choices under fixed budgets in biochemical optimization problems.

Summary

Researchers benchmarked five frontier language models on seven combinatorial datasets spanning protein engineering, reaction optimization, molecular design, peptide self-assembly and catalysis. Chemical context helped on average but with high variance, and no tested configuration decisively beat a mean statistical baseline across domains. Belief and action diagnostics showed that models intended to explore but behaved exploitatively when recent history remained in context.

AI role

Five frontier models selected experiments in Bayesian-optimization loops using combinations of literature priors, chemical context and accumulating observations.

Narrative role

Qualifies autonomous-lab claims with a cross-domain failure mode: verbal scientific intent and access to prior knowledge did not reliably translate into effective exploration or stronger experimental choices.

Caveat

This is a preprint evaluation over historical datasets and simulated sequential choices, not a prospective physical laboratory campaign. Results may depend on prompting, memory design and the selected statistical baseline.