Kevin-32B
Kevin-32B: Multi-Turn RL for Writing CUDA KernelsAn open-weights model built on QwQ-32B and trained with multi-turn reinforcement learning to write CUDA kernels for KernelBench tasks and refine them over several turns using execution feedback.
QwQ-32B is trained by RL to write CUDA kernels for deep learning operations, rewarded on measured correctness and speedup, and it refines its own kernels across turns using compiler and runtime feedback. Faster kernels for neural network operations lower the cost of training and serving models, and the RL recipe turns a kernel benchmark into a training signal for the next kernel-writing model. The authors had to block reward hacks such as returning the PyTorch reference instead of a kernel.
- QwQ-32B writes CUDA kernels
- Model refines kernels using compiler feedback
- RL rewards model on measured speedup
- QwQ-32B writes CUDA kernels
- Model refines kernels using compiler feedback
- RL rewards model on measured speedup
Why it is a road to recursion
It trains a model on the speed of the code that runs models, a reward that can be reused to train each next kernel writer, which is one concrete route for AI to cut its own compute bill.
Evidence
Per the paper, multi-turn RL raised kernel correctness from 56% for the QwQ-32B base model to 82% and mean speedup over the baseline from 0.53x to 1.10x.
Related loops
More kernels and compilers →- Gemini models propose code changes
- Evaluators keep best scoring versions
- Finds better Gemini training kernels
- Lowers cost to train next Gemini models
- repeat
- LLM agents write and benchmark kernels
- Kernels speed up AI model training
- Session data post-trains smaller kernel models
- repeat
- Coding agent edits LLM training code
- Agent benchmarks code on TPU hardware
- Verdict becomes context for next hypothesis
- Wins raise training throughput for same models
- repeat