Capability benchmark · Biology · Medicine
GeneBench measures multi-stage genomics and quantitative biology inference
A benchmark for realistic multi-stage scientific data analysis in genomics and quantitative biology.
Summary
OpenAI introduced GeneBench to test agents on multi-stage scientific analysis in genomics and adjacent quantitative biology domains. The benchmark uses 103 verifiable evaluations with realistic data-quality, model-selection, and diagnostic decision points, and reports that even the strongest tested settings remain far from saturation.
AI role
Agents explored staged datasets, performed quality control and statistical analysis, made chained inference decisions, and returned verifiable quantitative answers.
Narrative role
This adds a high-signal frontier-lab benchmark focused on the day-to-day analytical bottlenecks of data-rich biological research, with clear evidence of both progress and unreliability.
Caveat
The source is a frontier-lab benchmark paper, not peer reviewed, and several evaluated GPT results use internal or Pro harness settings that are not directly comparable to all external systems.