Alphabell.
Benchmark · AI R&D evals

LLM Speedrunning Benchmark

The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements19 tasks built from the NanoGPT speedrun, a community competition to train a GPT-2 model as fast as possible: the agent gets the previous record's training script, optionally with hints, and must reproduce the next record's speedup.

The loop

An LLM agent rewrites a GPT-2 training script to make LLM training faster, and the benchmark scores what fraction of the human record's speedup it recovers. The changes are the same kind (algorithmic and hardware-aware) that make training of the agent's own model family cheaper, so success is a direct step toward agents that speed up their own training.

The loop
LLM Speedrunning Benchmark
Benchmark · arXiv
  1. Agent rewrites LLM training script
  2. Script makes LLM training faster
  3. Benchmark scores the recovered speedup
  4. Speeds up agent's own model training
  1. Agent rewrites LLM training script
  2. Script makes LLM training faster
  3. Benchmark scores the recovered speedup
  4. Speeds up agent's own model training
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

Training-speed gains compound, so an agent that can find them makes every later training run of its successors cheaper.

Evidence

With pseudocode hints, o3-mini in a search scaffold recovered about 40% of the human speedup on average, and about 46% with pseudocode combined with text and mini-paper hints.

benchmarknanogpt-speedruntraining-efficiencyreproductionagents