KernelBook + KernelLLM
KernelBook and KernelLLMKernelBook is a dataset of 18,162 PyTorch programs from open-source GitHub code paired with equivalent Triton kernels generated by torch.compile. Meta's KernelLLM is an 8B model fine-tuned from Llama 3.1 8B Instruct on about 25,000 such PyTorch-to-Triton pairs to translate PyTorch modules into Triton GPU kernels.
KernelLLM, an 8B Llama 3.1 model fine-tuned on KernelBook's PyTorch-to-Triton pairs, writes Triton GPU kernels for the PyTorch modules that neural networks are built from. Faster kernels make training and inference cheaper, and kernels a model writes that pass verification can be added back to the corpus as training data for the next kernel model, the step that would close the loop.
- KernelLLM writes Triton GPU kernels
- Faster kernels make AI training cheaper
- Verified kernels train next kernel model
- KernelLLM writes Triton GPU kernels
- Faster kernels make AI training cheaper
- Verified kernels train next kernel model
Why it is a road to recursion
An open, verified corpus of kernel pairs lets small models learn to write the kernels that make larger models cheaper to run, and it can grow with each verified generation of model output.
Evidence
On KernelBench-Triton Level 1, KernelLLM scores 20.2 pass@1, above GPT-4o (15) and DeepSeek V3 (16), per its model card.
Related loops
More kernels and compilers →- Gemini models propose code changes
- Evaluators keep best scoring versions
- Finds better Gemini training kernels
- Lowers cost to train next Gemini models
- repeat
- LLM agents write and benchmark kernels
- Kernels speed up AI model training
- Session data post-trains smaller kernel models
- repeat
- Coding agent edits LLM training code
- Agent benchmarks code on TPU hardware
- Verdict becomes context for next hypothesis
- Wins raise training throughput for same models
- repeat