MLE-bench
MLE-bench: Evaluating Machine Learning Agents on Machine Learning EngineeringA benchmark of 75 Kaggle competitions that measures how well AI agents do end-to-end ML engineering (preparing data, training models, running experiments), graded against human Kaggle leaderboards.
AI agents are given a Kaggle task and must build, train and submit an ML model, and their submissions are scored against human medal thresholds. The paper states the loop it is tracking: agents able to do open-ended ML research at the level of improving their own training code could improve frontier models much faster than human researchers.
- Agents build and train ML models
- Submissions scored against human thresholds
- Measures open-ended ML research ability
- Agents improve their own training code
- Agents build and train ML models
- Submissions scored against human thresholds
- Measures open-ended ML research ability
- Agents improve their own training code
Why it is a road to recursion
The paper ties it to the model-autonomy category of OpenAI's Preparedness Framework, as a measure of progress toward agents that improve their own training code.
Evidence
At release the best setup (o1-preview with AIDE scaffolding) reached at least a Kaggle bronze medal in 16.9% of competitions, and the repo leaderboard's top entry (Famou-Agent 2.0 on Gemini-3-Pro-Preview, Feb 2026) lists 64.44% any-medal across all competitions.
Related loops
More ai r&d evals →- Coding agent writes fine-tuning code
- Agent post-trains a base LLM
- Trained model scored on target benchmark
- One turn of model training another
- repeat
- Agent rewrites LLM training script
- Script makes LLM training faster
- Benchmark scores the recovered speedup
- Speeds up agent's own model training
- repeat
- Agent replicates published ML research
- LLM judge grades the replication
- Measures ability to do ML research
- Research produces better future models
- repeat