autoresearch
A small open-source setup in which an AI coding agent edits a single-GPU LLM training script (a simplified nanochat), trains for a fixed 5 minutes, and keeps or discards each change based on validation bits per byte.
A coding agent edits train.py (architecture, optimizer, hyperparameters, batch size), runs a 5-minute training job, and keeps the change only if validation bits per byte improves, at about 12 experiments an hour. Karpathy ran it on nanochat: changes found by Claude running autonomously for about 2 days on a small model were merged into the nanochat training code, transferred to larger models, and became the base for a second round, with each round recorded on nanochat's Time-to-GPT-2 leaderboard.
- Coding agent edits LLM training script
- Runs a 5-minute training job
- Keeps change if validation improves
- Changes transferred to larger models
- Coding agent edits LLM training script
- Runs a 5-minute training job
- Keeps change if validation improves
- Changes transferred to larger models
Why it is a road to recursion
It is a working small-scale case of an AI agent improving a language model's training code with the gains kept and stacked, the same loop that becomes recursive once the agent's own model is the one being trained.
Evidence
nanochat's leaderboard records 'autoresearch round 1' cutting Time to GPT-2 from 2.02 to 1.80 hours (Mar 9 2026) and 'autoresearch round 2' to 1.65 hours (Mar 14 2026), and the round-1 commit says the changes were developed by Claude running autonomously over about 2 days.
Related loops
More research automation →- Qwen3-14B edits GPT training script
- Edit is trained and scored
- Score rewards Qwen3-14B weights update
- Trained checkpoint re-run in autoresearch loop
- repeat
- LLM agents propose hypotheses on AI tasks
- Agents implement and test them
- Findings Memory steers later proposals
- Validated finding directly improves AI system
- repeat
- LLM agents propose ML research ideas
- Agents run experiments with tree search
- Agents write up results and papers
- Output is new knowledge about models
- repeat