Alphabell.
Benchmark · AI R&D evals

MLE-bench

MLE-bench: Evaluating Machine Learning Agents on Machine Learning EngineeringA benchmark of 75 Kaggle competitions that measures how well AI agents do end-to-end ML engineering (preparing data, training models, running experiments), graded against human Kaggle leaderboards.

The loop

AI agents are given a Kaggle task and must build, train and submit an ML model, and their submissions are scored against human medal thresholds. The paper states the loop it is tracking: agents able to do open-ended ML research at the level of improving their own training code could improve frontier models much faster than human researchers.

The loop
MLE-bench
Benchmark · GitHub
  1. Agents build and train ML models
  2. Submissions scored against human thresholds
  3. Measures open-ended ML research ability
  4. Agents improve their own training code
  1. Agents build and train ML models
  2. Submissions scored against human thresholds
  3. Measures open-ended ML research ability
  4. Agents improve their own training code
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

The paper ties it to the model-autonomy category of OpenAI's Preparedness Framework, as a measure of progress toward agents that improve their own training code.

Evidence

At release the best setup (o1-preview with AIDE scaffolding) reached at least a Kaggle bronze medal in 16.9% of competitions, and the repo leaderboard's top entry (Famou-Agent 2.0 on Gemini-3-Pro-Preview, Feb 2026) lists 64.44% any-medal across all competitions.

benchmarkml-engineeringkaggleagentsopenai