Alphabell.
Dataset · Kernels and compilers

KernelBook + KernelLLM

KernelBook and KernelLLMKernelBook is a dataset of 18,162 PyTorch programs from open-source GitHub code paired with equivalent Triton kernels generated by torch.compile. Meta's KernelLLM is an 8B model fine-tuned from Llama 3.1 8B Instruct on about 25,000 such PyTorch-to-Triton pairs to translate PyTorch modules into Triton GPU kernels.

The loop

KernelLLM, an 8B Llama 3.1 model fine-tuned on KernelBook's PyTorch-to-Triton pairs, writes Triton GPU kernels for the PyTorch modules that neural networks are built from. Faster kernels make training and inference cheaper, and kernels a model writes that pass verification can be added back to the corpus as training data for the next kernel model, the step that would close the loop.

The loop
KernelBook + KernelLLM
Dataset · Hugging Face
  1. KernelLLM writes Triton GPU kernels
  2. Faster kernels make AI training cheaper
  3. Verified kernels train next kernel model
  1. KernelLLM writes Triton GPU kernels
  2. Faster kernels make AI training cheaper
  3. Verified kernels train next kernel model
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

An open, verified corpus of kernel pairs lets small models learn to write the kernels that make larger models cheaper to run, and it can grow with each verified generation of model output.

Evidence

On KernelBench-Triton Level 1, KernelLLM scores 20.2 pass@1, above GPT-4o (15) and DeepSeek V3 (16), per its model card.

tritondatasetfine-tuningpytorchgpu-kernels