Capability benchmark · Biology · Computer science
BixBench evaluates agents on open-ended computational-biology analysis
A bioinformatics benchmark with more than 50 real-world scenarios and nearly 300 open-answer research questions.
Summary
BixBench, announced in 2025 by FutureHouse and partners, benchmarks LLM-based agents on open-ended bioinformatics and computational-biology analysis scenarios. The paper reports low frontier-model performance, including 17% open-answer accuracy, despite a custom agent framework.
AI role
Tested LLM-based agents on exploring biological datasets, running long multi-step analyses, and interpreting nuanced computational-biology results.
Narrative role
This historical-gap backfill adds an earlier realistic biology-agent benchmark that helps anchor the 2025 shift from knowledge tests toward open-ended computational research workflows.
Caveat
The benchmark is preprint-status and was built by an organization developing research agents, so its task selection and agent implementation should be interpreted with that context.