Alphabell.
Benchmark · AI R&D evals

PaperBench

PaperBench: Evaluating AI's Ability to Replicate AI ResearchA benchmark in which AI agents must replicate 20 ICML 2024 papers from scratch: understand the paper, build the codebase and run the experiments. Attempts are graded by an LLM judge against 8,316 rubric items co-written with the papers' authors.

The loop

An LLM agent reads a published ML paper, rebuilds its code and reruns its experiments, and an LLM judge grades the replication against author-written rubrics. The score measures how close agents are to carrying out the research that produces better models, and the paper ties this capability to the model autonomy and ML R&D thresholds in OpenAI, Anthropic and Google DeepMind safety frameworks.

The loop
PaperBench
Benchmark · arXiv
  1. Agent replicates published ML research
  2. LLM judge grades the replication
  3. Measures ability to do ML research
  4. Research produces better future models
  1. Agent replicates published ML research
  2. LLM judge grades the replication
  3. Measures ability to do ML research
  4. Research produces better future models
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

Reproducing published ML results end to end is a precondition for an agent that improves its own training stack, so this score tracks how far the loop is from running without researchers.

Evidence

The best agent, Claude 3.5 Sonnet (New) with an open-source scaffold, averaged a 21.0% replication score, while ML PhDs reached 41.4% after 48 hours on a 3-paper subset versus 26.6% for o1.

benchmarkpaper-replicationml-researchagentsllm-judge