Capability benchmark · Biology · Computer science

LABBench2 raises the bar for biology research-agent evaluation

A more realistic biology research benchmark with nearly 1,900 tasks across practical scientific work contexts.

Summary

LABBench2 extends the earlier Language Agent Biology Benchmark with nearly 1,900 tasks designed to test useful biology research capabilities in more realistic contexts. The authors report that frontier models improved on the older benchmark but face a meaningful difficulty increase on LABBench2.

AI role

Evaluated AI systems on biology research tasks such as literature use, figures, protocols, molecular biology, and tool-supported workflows.

Narrative role

This is a benchmark follow-up that clarifies whether apparent progress on biology research-agent tasks survives a more realistic evaluation design.

Caveat

The benchmark is a preprint and mostly continues the LAB-Bench task family; it measures capability proxies rather than direct scientific discoveries.

Related events