Capability benchmark · Computer science · General science
AutoResearchBench tests agents on scientific literature discovery
A benchmark for deep and wide scientific literature discovery tasks with open-ended search requirements.
Summary
AutoResearchBench evaluates whether AI agents can perform scientific literature discovery through two tasks: finding specific target papers and collecting sets of papers that satisfy open-ended scientific conditions. The reported results indicate that strong models still perform poorly on these research-oriented search tasks.
AI role
Agents search for target papers or collect qualifying scientific papers by reasoning over concepts, evidence, and detailed literature constraints.
Narrative role
Literature discovery is a core research workflow, so this benchmark helps separate broad web-browsing competence from science-specific evidence gathering.
Caveat
The source is a preprint and evaluates literature search rather than direct hypothesis testing or experimental discovery.