Capability benchmark · Biology · Computer science

BixBench3 scales biology-agent evaluation to complete computational studies

A 20-study benchmark with 67 GB average raw inputs and 138 graded artifacts for end-to-end computational-biology analysis.

Summary

Edison Scientific's BixBench3 evaluates whether agents can reconstruct complete computational-biology studies from raw data and a prescribed high-level analysis plan. Thirteen frontier models attempted 20 studies and 138 graded artifacts; the top average score was 0.48. Performance fell on larger datasets and deeper analysis chains, while runs averaged 6.8 hours and 102 million tokens.

AI role

Agents received a research objective, high-level methodological guidance, and raw biological data, then built and executed complete analysis pipelines whose artifacts were graded against published results.

Narrative role

BixBench3 marks a meaningful step beyond the original BixBench's open-ended questions and isolated analyses: it measures long-horizon execution at the scale, data volume, and dependency depth of full computational studies.

Caveat

The benchmark prescribes the research question and analysis plan, so it tests execution rather than scientific judgment about which questions or methods to pursue. Deterministic grading can penalize valid alternative analyses, inherits errors from source papers, and is reported in an Edison-authored preprint without independent replication.

Related events