Capability benchmark · Computer science

ML research agents explore novel solutions but rarely convert novelty into performance

A human-calibrated evaluation of novelty and usefulness across multi-turn agent trajectories on ten machine-learning competitions.

Summary

The study evaluated two research-agent frameworks and two model families on ten Kaggle-derived MLE-Bench tasks. Agents often entered solution regions judged more novel than medal-winning human notebooks, but novelty declined as runs shifted from exploration to refinement and did not predict score improvement. GPT-5 exceeded the human historical-novelty baseline on nine of ten tasks, yet only 21.25% of its runs reached medal-level performance.

AI role

AIDE and AIRA-Dojo agents iteratively proposed and tested ML solutions; separate models summarized trajectories and judged novelty against prior agent attempts and large human notebook corpora.

Narrative role

Separates a plausible ingredient of discovery—exploring unfamiliar approaches—from usefulness. The results caution against interpreting novelty or divergent search alone as scientific progress when agents cannot reliably turn it into better outcomes.

Caveat

The tasks are Kaggle-style ML engineering rather than open-ended science. Historical novelty depends on retrieval and LLM judgment over incomplete public notebook corpora, although the P-creativity judge was calibrated on 300 human-annotated episodes.