Capability benchmark · Computer science · General science

Autoresearch Bench measures four-hour agent experimentation loops

A benchmark of agents that iteratively modify solutions, run experiments and use external graders across optimization, training, inference and scientific-machine-learning tasks.

Summary

Autoresearch Bench evaluates Claude Opus 5, GPT-5.6 Sol, Gemini 3.7 Flash, Kimi K3 and Grok 4.6 on iterative optimization and scientific-computing tasks. Agents receive fixed four-hour budgets and can repeatedly test and grade their own changes; most tasks include a private split. Claude Opus 5 ranked first, showed roughly twice the relative performance of peers on optimization tasks and generalized best on two of three tasks with notable public-private score disagreement.

AI role

Five frontier models independently ran four-hour research loops in isolated workspaces, choosing experiments and submitting candidate solutions to graders with private held-out data.

Narrative role

Adds a controlled measure of experiment-driven research behavior rather than static answers or one-shot code generation, including held-out evaluation and trajectories that reveal overfitting, submission strategy and early stopping.

Caveat

This is a first-party benchmark report rather than a peer-reviewed study. Cross-task normalized scores use a frozen average as divisor, failed runs are excluded from aggregate means, and bounded coding and optimization tasks remain proxies for real research.