Alphabell.
Mini-project · Kernels and compilers

Stanford AI-generated kernels

Surprisingly Fast AI-Generated Kernels We Didn't Mean to Publish (Yet)A report on a test-time search in which o3 and Gemini 2.5 Pro first write optimization ideas in natural language and then branch into many kernel implementations per round, run for 5 rounds on 10 KernelBench Level 1 problems. Several generated FP32 kernels matched or beat PyTorch on an L40S GPU.

The loop

Frontier LLMs (o3, Gemini 2.5 Pro) reason in language about optimizations, write CUDA kernels for core ML operations, and seed each new round with the fastest variants. The authors started the work to generate synthetic data for training better kernel-generation models, so the search output is meant to feed the next model that writes kernels for AI workloads.

The loop
Stanford AI-generated kernels
Mini-project · Stanford CRFM
  1. Frontier LLMs write CUDA kernels
  2. Fastest variants seed each new round
  3. Output feeds next kernel-writing model
  1. Frontier LLMs write CUDA kernels
  2. Fastest variants seed each new round
  3. Output feeds next kernel-writing model
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

It ties test-time search to training data: kernels that frontier models find by search are intended to train the next kernel-writing model, a direct data loop for AI performance engineering.

Evidence

On an L40S, generated kernels reached 484.4% of PyTorch performance for LayerNorm, 179.9% for Conv2D and 101.3% for FP32 matmul, though the authors note FP32 is less common in modern ML and their FP16 matmul and FlashAttention kernels still lagged.

cudatest-time-searchsynthetic-datakernelbenchllm-reasoning