Capability benchmark · Computer science · Biology
LAB-Bench evaluates language agents on practical biology research tasks
Practical biology research-task benchmark covering literature, figures, databases, and sequences.
Summary
The arXiv paper introduces LAB-Bench, a benchmark of more than 2,400 questions for evaluating AI systems on practical biology research capabilities, including literature reasoning, figure interpretation, database navigation and DNA or protein sequence manipulation.
AI role
Performed tasks closer to biology research assistance than ordinary science exam questions.
Narrative role
LAB-Bench helps the tracker separate biology-agent capability from generic chatbot ability.
Caveat
Benchmark questions still simplify real laboratory planning, execution, and interpretation.