Capability benchmark · Biology · Medicine · Computer science

BAISBench evaluates AI scientists on real single-cell discovery tasks

A peer-reviewed benchmark pairing 15 expert-labeled single-cell datasets with 193 data-driven discovery questions derived from 41 published single-cell studies.

Summary

BAISBench evaluates AI scientist systems on real single-cell transcriptomic analysis rather than static biology questions. Its two tasks test cell-type annotation and recovery of conclusions from published single-cell studies; the study finds systems can approach graduate-level performance on some discovery-oriented tasks but retain deficits where biological judgment and interpretation dominate.

AI role

AI scientist systems process single-cell data, annotate cell types, and analyze datasets to identify biological conclusions consistent with published findings.

Narrative role

This adds peer-reviewed, data-grounded evidence about a key boundary for AI-assisted biology: systems can perform parts of a realistic omics workflow, while human expertise remains important for nuanced biological interpretation.

Caveat

The discovery task is framed as multiple-choice recovery of established findings from curated single-cell studies, not generation and experimental validation of genuinely new biological discoveries; results may not transfer to wet-lab or other omics workflows.

Related events