Capability benchmark · Biology · Computer science

HeurekaBench grounds co-scientist evaluation in data-backed single-cell workflows

A benchmark framework and single-cell biology instance for open-ended, dataset-grounded scientific agent questions.

Summary

HeurekaBench proposes a framework for constructing benchmark questions from scientific studies, code repositories, and experimental datasets. Its single-cell biology instance tests whether agents can design analyses and produce data-backed answers to open-ended research questions.

AI role

Evaluated agents on multi-step analysis of experimental datasets, requiring data-backed responses rather than factual recall alone.

Narrative role

This adds a capability benchmark focused on the open-ended co-scientist role, where the measured task is exploratory analysis rather than static question answering.

Caveat

The initial instantiation is narrow to single-cell biology and uses LLM-as-judge evaluation for free-form outputs, so results depend on benchmark construction and grading choices.