Capability benchmark · Biology · Computer science

BixBench evaluates agents on open-ended computational-biology analysis

A bioinformatics benchmark with more than 50 real-world scenarios and nearly 300 open-answer research questions.

Summary

BixBench, announced in 2025 by FutureHouse and partners, benchmarks LLM-based agents on open-ended bioinformatics and computational-biology analysis scenarios. The paper reports low frontier-model performance, including 17% open-answer accuracy, despite a custom agent framework.

AI role

Tested LLM-based agents on exploring biological datasets, running long multi-step analyses, and interpreting nuanced computational-biology results.

Narrative role

This historical-gap backfill adds an earlier realistic biology-agent benchmark that helps anchor the 2025 shift from knowledge tests toward open-ended computational research workflows.

Caveat

The benchmark is preprint-status and was built by an organization developing research agents, so its task selection and agent implementation should be interpreted with that context.