Capability benchmark · Biology · Computer science
LABBench2 raises the bar for biology research-agent evaluation
A more realistic biology research benchmark with nearly 1,900 tasks across practical scientific work contexts.
Summary
LABBench2 extends the earlier Language Agent Biology Benchmark with nearly 1,900 tasks designed to test useful biology research capabilities in more realistic contexts. The authors report that frontier models improved on the older benchmark but face a meaningful difficulty increase on LABBench2.
AI role
Evaluated AI systems on biology research tasks such as literature use, figures, protocols, molecular biology, and tool-supported workflows.
Narrative role
This is a benchmark follow-up that clarifies whether apparent progress on biology research-agent tasks survives a more realistic evaluation design.
Caveat
The benchmark is a preprint and mostly continues the LAB-Bench task family; it measures capability proxies rather than direct scientific discoveries.