Capability benchmark · Computer science · Mathematics · Biology · General science
AIRS-Bench evaluates agents across the machine-learning research lifecycle
A 20-task benchmark for AI research agents using tasks sourced from state-of-the-art machine-learning papers.
Summary
AIRS-Bench introduces a benchmark for frontier AI research science agents across parts of the research lifecycle, including idea generation, experiment analysis, and iterative refinement. The reported baselines show mixed capability: agents beat human state-of-the-art on a few tasks but remain below it on most tasks.
AI role
Agents generate ideas, analyze experiments, and iteratively refine solutions without being given baseline code.
Narrative role
This adds a current capability benchmark that directly targets research-agent behavior instead of only question answering or coding snippets.
Caveat
The paper is a preprint and the benchmark is centered on machine-learning research tasks, so it is not a general measure of all scientific work.