Capability benchmark · Computer science · General science
Autoresearch Bench measures four-hour agent experimentation loops
A benchmark of agents that iteratively modify solutions, run experiments and use external graders across optimization, training, inference and scientific-machine-learning tasks.
Summary
Autoresearch Bench evaluates Claude Opus 5, GPT-5.6 Sol, Gemini 3.7 Flash, Kimi K3 and Grok 4.6 on iterative optimization and scientific-computing tasks. Agents receive fixed four-hour budgets and can repeatedly test and grade their own changes; most tasks include a private split. Claude Opus 5 ranked first, showed roughly twice the relative performance of peers on optimization tasks and generalized best on two of three tasks with notable public-private score disagreement.
AI role
Five frontier models independently ran four-hour research loops in isolated workspaces, choosing experiments and submitting candidate solutions to graders with private held-out data.
Narrative role
Adds a controlled measure of experiment-driven research behavior rather than static answers or one-shot code generation, including held-out evaluation and trajectories that reveal overfitting, submission strategy and early stopping.
Caveat
This is a first-party benchmark report rather than a peer-reviewed study. Cross-task normalized scores use a frozen average as divisor, failed runs are excluded from aggregate means, and bounded coding and optimization tasks remain proxies for real research.