Capability benchmark · Computer science · General science

ScienceAgentBench tests agents on data-driven discovery tasks

Agent performance on research tasks extracted from peer-reviewed data-driven science papers.

Summary

ScienceAgentBench reframes data-driven scientific discovery problems as executable tasks and reports that even the best tested agents solved only a minority independently.

AI role

Generated code and analyses for realistic data-driven scientific discovery tasks.

Narrative role

This supports the timeline by measuring capability gaps between demos and realistic scientific data work.

Caveat

Benchmark performance is an upstream indicator; it is not evidence of field-level acceleration by itself.