Capability benchmark · Computer science
ML research agents explore novel solutions but rarely convert novelty into performance
A human-calibrated evaluation of novelty and usefulness across multi-turn agent trajectories on ten machine-learning competitions.
Summary
The study evaluated two research-agent frameworks and two model families on ten Kaggle-derived MLE-Bench tasks. Agents often entered solution regions judged more novel than medal-winning human notebooks, but novelty declined as runs shifted from exploration to refinement and did not predict score improvement. GPT-5 exceeded the human historical-novelty baseline on nine of ten tasks, yet only 21.25% of its runs reached medal-level performance.
AI role
AIDE and AIRA-Dojo agents iteratively proposed and tested ML solutions; separate models summarized trajectories and judged novelty against prior agent attempts and large human notebook corpora.
Narrative role
Separates a plausible ingredient of discovery—exploring unfamiliar approaches—from usefulness. The results caution against interpreting novelty or divergent search alone as scientific progress when agents cannot reliably turn it into better outcomes.
Caveat
The tasks are Kaggle-style ML engineering rather than open-ended science. Historical novelty depends on retrieval and LLM judgment over incomplete public notebook corpora, although the P-creativity judge was calibrated on 300 human-annotated episodes.