Capability benchmark · Computer science
MLE-bench evaluates agents on Kaggle-style machine-learning engineering
Agent performance on machine-learning engineering tasks from Kaggle competitions.
Summary
The arXiv paper introduces MLE-bench, a benchmark built from 75 Kaggle competitions to test AI agents on machine-learning engineering. Agents must inspect data, write and run training code, iterate on experiments and submit solutions under competition-like scoring.
AI role
Inspected data, wrote training code, iterated on experiments, and submitted scored solutions.
Narrative role
MLE-bench is a supporting signal for AI research acceleration because ML experimentation is a core research labor bottleneck.
Caveat
Kaggle-style performance is not the same as open-ended scientific question selection or interpretation.