Capability benchmark · Computer science · General science
ScienceAgentBench tests agents on data-driven discovery tasks
Agent performance on research tasks extracted from peer-reviewed data-driven science papers.
Summary
ScienceAgentBench reframes data-driven scientific discovery problems as executable tasks and reports that even the best tested agents solved only a minority independently.
AI role
Generated code and analyses for realistic data-driven scientific discovery tasks.
Narrative role
This supports the timeline by measuring capability gaps between demos and realistic scientific data work.
Caveat
Benchmark performance is an upstream indicator; it is not evidence of field-level acceleration by itself.