Alphabell.
Benchmark · AI R&D evals

RE-Bench

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human expertsSeven open-ended ML research-engineering environments (such as optimizing a GPU kernel, fixing corrupted embeddings, fitting a scaling law and finetuning GPT-2 with RL), with baselines from 71 eight-hour attempts by 61 human experts.

The loop

Language-model agents are placed in realistic AI R&D tasks, such as speeding up an LLM finetuning script or a GPU kernel, and scored against human experts at matched time budgets. METR built it because frontier AI safety policies flag automation of AI R&D by AI agents as a key risk, so it measures how close agents are to doing the work that improves AI systems.

The loop
RE-Bench
Benchmark · METR
  1. Agents do AI research tasks
  2. Scored against human expert baselines
  3. Measures automation of AI research
  4. Research improves future AI systems
  1. Agents do AI research tasks
  2. Scored against human expert baselines
  3. Measures automation of AI research
  4. Research improves future AI systems
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

It directly measures, against human experts, how far agents are from automating the AI research engineering that would let AI accelerate its own development.

Evidence

The best AI agents scored 4x higher than human experts with a 2-hour budget per environment, while humans narrowly exceeded the top agents at 8 hours and reached 2x their score at 32 hours.

benchmarkai-rndhuman-baselinemetrsafety