Alphabell.
Project · Kernels and compilers

Kevin-32B

Kevin-32B: Multi-Turn RL for Writing CUDA KernelsAn open-weights model built on QwQ-32B and trained with multi-turn reinforcement learning to write CUDA kernels for KernelBench tasks and refine them over several turns using execution feedback.

The loop

QwQ-32B is trained by RL to write CUDA kernels for deep learning operations, rewarded on measured correctness and speedup, and it refines its own kernels across turns using compiler and runtime feedback. Faster kernels for neural network operations lower the cost of training and serving models, and the RL recipe turns a kernel benchmark into a training signal for the next kernel-writing model. The authors had to block reward hacks such as returning the PyTorch reference instead of a kernel.

The loop
Kevin-32B
Project · Cognition
  1. QwQ-32B writes CUDA kernels
  2. Model refines kernels using compiler feedback
  3. RL rewards model on measured speedup
  1. QwQ-32B writes CUDA kernels
  2. Model refines kernels using compiler feedback
  3. RL rewards model on measured speedup
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

It trains a model on the speed of the code that runs models, a reward that can be reused to train each next kernel writer, which is one concrete route for AI to cut its own compute bill.

Evidence

Per the paper, multi-turn RL raised kernel correctness from 56% for the QwQ-32B base model to 82% and mean speedup over the baseline from 0.53x to 1.10x.

cudareinforcement-learningmulti-turnkernelbenchopen-weights