Alphabell.
Mini-project · Research automation

autoresearch

A small open-source setup in which an AI coding agent edits a single-GPU LLM training script (a simplified nanochat), trains for a fixed 5 minutes, and keeps or discards each change based on validation bits per byte.

The loop

A coding agent edits train.py (architecture, optimizer, hyperparameters, batch size), runs a 5-minute training job, and keeps the change only if validation bits per byte improves, at about 12 experiments an hour. Karpathy ran it on nanochat: changes found by Claude running autonomously for about 2 days on a small model were merged into the nanochat training code, transferred to larger models, and became the base for a second round, with each round recorded on nanochat's Time-to-GPT-2 leaderboard.

The loop
autoresearch
Mini-project · GitHub
  1. Coding agent edits LLM training script
  2. Runs a 5-minute training job
  3. Keeps change if validation improves
  4. Changes transferred to larger models
  1. Coding agent edits LLM training script
  2. Runs a 5-minute training job
  3. Keeps change if validation improves
  4. Changes transferred to larger models
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

It is a working small-scale case of an AI agent improving a language model's training code with the gains kept and stacked, the same loop that becomes recursive once the agent's own model is the one being trained.

Evidence

nanochat's leaderboard records 'autoresearch round 1' cutting Time to GPT-2 from 2.02 to 1.80 hours (Mar 9 2026) and 'autoresearch round 2' to 1.65 hours (Mar 14 2026), and the round-1 commit says the changes were developed by Claude running autonomously over about 2 days.

coding-agentllm-trainingnanochathyperparameter-searchhobbyist