PaperBench
PaperBench: Evaluating AI's Ability to Replicate AI ResearchA benchmark in which AI agents must replicate 20 ICML 2024 papers from scratch: understand the paper, build the codebase and run the experiments. Attempts are graded by an LLM judge against 8,316 rubric items co-written with the papers' authors.
An LLM agent reads a published ML paper, rebuilds its code and reruns its experiments, and an LLM judge grades the replication against author-written rubrics. The score measures how close agents are to carrying out the research that produces better models, and the paper ties this capability to the model autonomy and ML R&D thresholds in OpenAI, Anthropic and Google DeepMind safety frameworks.
- Agent replicates published ML research
- LLM judge grades the replication
- Measures ability to do ML research
- Research produces better future models
- Agent replicates published ML research
- LLM judge grades the replication
- Measures ability to do ML research
- Research produces better future models
Why it is a road to recursion
Reproducing published ML results end to end is a precondition for an agent that improves its own training stack, so this score tracks how far the loop is from running without researchers.
Evidence
The best agent, Claude 3.5 Sonnet (New) with an open-source scaffold, averaged a 21.0% replication score, while ML PhDs reached 41.4% after 48 hours on a 3-paper subset versus 26.6% for o1.
Related loops
More ai r&d evals →- Coding agent writes fine-tuning code
- Agent post-trains a base LLM
- Trained model scored on target benchmark
- One turn of model training another
- repeat
- Agent rewrites LLM training script
- Script makes LLM training faster
- Benchmark scores the recovered speedup
- Speeds up agent's own model training
- repeat
- LLM writes replacement GPU kernel
- Kernel is timed against baseline
- Workloads are neural network operators
- Makes agent's own training faster
- repeat