Capability benchmark · Computer science

MLE-bench evaluates agents on Kaggle-style machine-learning engineering

Agent performance on machine-learning engineering tasks from Kaggle competitions.

Summary

The arXiv paper introduces MLE-bench, a benchmark built from 75 Kaggle competitions to test AI agents on machine-learning engineering. Agents must inspect data, write and run training code, iterate on experiments and submit solutions under competition-like scoring.

AI role

Inspected data, wrote training code, iterated on experiments, and submitted scored solutions.

Narrative role

MLE-bench is a supporting signal for AI research acceleration because ML experimentation is a core research labor bottleneck.

Caveat

Kaggle-style performance is not the same as open-ended scientific question selection or interpretation.