Capability benchmark · Computer science · Biology

LAB-Bench evaluates language agents on practical biology research tasks

Practical biology research-task benchmark covering literature, figures, databases, and sequences.

Summary

The arXiv paper introduces LAB-Bench, a benchmark of more than 2,400 questions for evaluating AI systems on practical biology research capabilities, including literature reasoning, figure interpretation, database navigation and DNA or protein sequence manipulation.

AI role

Performed tasks closer to biology research assistance than ordinary science exam questions.

Narrative role

LAB-Bench helps the tracker separate biology-agent capability from generic chatbot ability.

Caveat

Benchmark questions still simplify real laboratory planning, execution, and interpretation.