Capability benchmark · Computer science · General science

AutoResearchBench tests agents on scientific literature discovery

A benchmark for deep and wide scientific literature discovery tasks with open-ended search requirements.

Summary

AutoResearchBench evaluates whether AI agents can perform scientific literature discovery through two tasks: finding specific target papers and collecting sets of papers that satisfy open-ended scientific conditions. The reported results indicate that strong models still perform poorly on these research-oriented search tasks.

AI role

Agents search for target papers or collect qualifying scientific papers by reasoning over concepts, evidence, and detailed literature constraints.

Narrative role

Literature discovery is a core research workflow, so this benchmark helps separate broad web-browsing competence from science-specific evidence gathering.

Caveat

The source is a preprint and evaluates literature search rather than direct hypothesis testing or experimental discovery.