TPU autoresearch
TPU Model Performance Auto-optimization (tpu_performance_autoresearch_wiki)A single-developer repo that specializes Karpathy's autoresearch loop to TPU performance engineering: a coding agent keeps a markdown wiki of profiles, framework source and past experiments, then hypothesizes, patches model code or writes Pallas kernels, benchmarks on real TPUs and keeps or discards each change. It has been run with Claude Code, Codex and Antigravity agents.
A coding agent (Codex, Claude and Gemini models in the case studies) reads profiler output and its own wiki, edits the training code of an LLM such as Qwen3-8B or authors a Pallas attention kernel, benchmarks it on TPU hardware and records the verdict as context for the next hypothesis. Wins raise training throughput (MFU) or kernel speed for LLMs, the same class of model doing the work, and kernel wins are merged only after they validate end to end in the model lane.
- Coding agent edits LLM training code
- Agent benchmarks code on TPU hardware
- Verdict becomes context for next hypothesis
- Wins raise training throughput for same models
- Coding agent edits LLM training code
- Agent benchmarks code on TPU hardware
- Verdict becomes context for next hypothesis
- Wins raise training throughput for same models
Why it is a road to recursion
Coding agents beating hand-tuned baselines at making LLM training and attention kernels faster on real hardware cut the cost of training the models behind the next, stronger agents.
Evidence
On Qwen3-8B training on a TPU v6e-8, a Codex GPT-5.5 agent reached 47.3% MFU at 2k context against 36.6% for hand-optimized MaxText, and in the kernel lane a Claude Opus 5 agent reached a 6.77x GQA-attention speedup over the naive baseline in 22 experiments, against 2.48x for MaxKernel's published best.
Related loops
More kernels and compilers →- Gemini models propose code changes
- Evaluators keep best scoring versions
- Finds better Gemini training kernels
- Lowers cost to train next Gemini models
- repeat
- LLM agents write and benchmark kernels
- Kernels speed up AI model training
- Session data post-trains smaller kernel models
- repeat
- Frontier LLMs write CUDA kernels
- Fastest variants seed each new round
- Output feeds next kernel-writing model
- repeat