Alphabell.
Benchmark · Inference efficiency

FlashInfer-Bench

FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM SystemsA benchmark and deployment framework for LLM-written GPU kernels in LLM serving: a shared trace schema, 41 kernel definitions with about 1,600 workloads taken from real serving traces, a correctness and performance harness, a public leaderboard, and an apply() hook that swaps the best kernel into SGLang or vLLM.

The loop

LLM agents (the paper evaluates GPT-5, Claude Opus 4.1, Gemini 2.5 Pro and o3) write GPU kernels for the operators that LLM serving engines run, such as attention, GEMM, normalization, sampling and MoE. The framework validates them on real serving workloads and substitutes the winners into production engines, so faster agent-written kernels make LLM inference cheaper, and new serving traces define the next round of kernel tasks.

The loop
FlashInfer-Bench
Benchmark · arXiv
  1. LLM agents write GPU kernels
  2. Faster kernels make LLM inference cheaper
  3. New serving traces define next kernel tasks
  1. LLM agents write GPU kernels
  2. Faster kernels make LLM inference cheaper
  3. New serving traces define next kernel tasks
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

It supplies the plumbing for a loop in which models write the kernels that serve models, with a standard path from a benchmark win to a production deployment.

Evidence

Agent-written kernels came close to or above human level on memory-bound RMSNorm, but Triton GEMM kernels reached under 50% of state-of-the-art performance on more than half of the workloads, and kernel substitution added under 0.8% end-to-end overhead.

llm-servinggpu-kernelsbenchmarksglangvllm