Alphabell.
Benchmark · AI R&D evals

KernelBench

KernelBench: Can LLMs Write Efficient GPU Kernels?A benchmark of 250 PyTorch ML workloads (single operators, fused operations and full model architectures) on which LLMs must write correct and faster GPU kernels, scored with the fast_p metric.

The loop

An LLM is given a PyTorch ML workload and writes a replacement GPU kernel, which is checked for correctness and timed against the PyTorch baseline, optionally with execution and profiling feedback for iterative refinement. Because the workloads are the operators and architectures used to train and run neural networks, a model that does well here can make ML training and inference, including its own, faster and cheaper.

The loop
KernelBench
Benchmark · GitHub
  1. LLM writes replacement GPU kernel
  2. Kernel is timed against baseline
  3. Workloads are neural network operators
  4. Makes agent's own training faster
  1. LLM writes replacement GPU kernel
  2. Kernel is timed against baseline
  3. Workloads are neural network operators
  4. Makes agent's own training faster
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

It measures the skill by which models could speed up the kernels in their own training and inference stack, an indirect but explicit path to AI making AI cheaper to build.

Evidence

The paper reports frontier reasoning models matched the PyTorch baseline in less than 20% of cases without refinement, with iterative execution and profiling feedback improving results.

benchmarkgpu-kernelscudaperformancestanford