Capability benchmark · Biology · Computer science
HeurekaBench grounds co-scientist evaluation in data-backed single-cell workflows
A benchmark framework and single-cell biology instance for open-ended, dataset-grounded scientific agent questions.
Summary
HeurekaBench proposes a framework for constructing benchmark questions from scientific studies, code repositories, and experimental datasets. Its single-cell biology instance tests whether agents can design analyses and produce data-backed answers to open-ended research questions.
AI role
Evaluated agents on multi-step analysis of experimental datasets, requiring data-backed responses rather than factual recall alone.
Narrative role
This adds a capability benchmark focused on the open-ended co-scientist role, where the measured task is exploratory analysis rather than static question answering.
Caveat
The initial instantiation is narrow to single-cell biology and uses LLM-as-judge evaluation for free-form outputs, so results depend on benchmark construction and grading choices.