Research and evals
AI R&D evals
Benchmarks that measure how well AI does AI research and engineering: the scoreboard for the loop.
PostTrainBench
- Coding agent writes fine-tuning code
- Agent post-trains a base LLM
- Trained model scored on target benchmark
- One turn of model training another
- repeat
AI R&D evalsELLIS Institute Tubingen, MPI-IS, University of Tubingen, Thoughtful Lab (Rank et al.) · 2026
LLM Speedrunning Benchmark
- Agent rewrites LLM training script
- Script makes LLM training faster
- Benchmark scores the recovered speedup
- Speeds up agent's own model training
- repeat
AI R&D evalsMeta and collaborators (Zhao et al.) · 2025
PaperBench
- Agent replicates published ML research
- LLM judge grades the replication
- Measures ability to do ML research
- Research produces better future models
- repeat
AI R&D evalsOpenAI · 2025
KernelBench
- LLM writes replacement GPU kernel
- Kernel is timed against baseline
- Workloads are neural network operators
- Makes agent's own training faster
- repeat
AI R&D evalsOuyang et al. (Stanford) · 2024
RE-Bench
- Agents do AI research tasks
- Scored against human expert baselines
- Measures automation of AI research
- Research improves future AI systems
- repeat
AI R&D evalsMETR · 2024
MLE-bench
- Agents build and train ML models
- Submissions scored against human thresholds
- Measures open-ended ML research ability
- Agents improve their own training code
- repeat
AI R&D evalsOpenAI · 2024