Alphabell.
Project · Kernels and compilers

KernelEvolve

KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at MetaMeta's agentic system that takes kernel specifications and generates and tunes kernels in Triton, CuTe DSL and lower-level languages for recommendation model training and inference on NVIDIA GPUs, AMD GPUs and Meta's MTIA accelerators, using graph-based search, retrieval-augmented prompts and a hardware knowledge base.

The loop

LLM agents write and benchmark kernels for Meta's production ranking models, and a job harness feeds measured performance and diagnostics back to the agent over hundreds of candidates. The kernels speed up training and inference of those AI models in production, and Meta says successful strategies are distilled into a reusable skill library and that session data is used to post-train smaller kernel models with RL rewarded by measured kernel performance.

The loop
KernelEvolve
Project · arXiv
  1. LLM agents write and benchmark kernels
  2. Kernels speed up AI model training
  3. Session data post-trains smaller kernel models
  1. LLM agents write and benchmark kernels
  2. Kernels speed up AI model training
  3. Session data post-trains smaller kernel models
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

It is a production case where agent-written kernels speed up a company's own model training and inference, and the agent's experience is recycled into skills and training data for the next kernel models.

Evidence

Meta reports over 60% inference throughput improvement for its Andromeda ads model on NVIDIA GPUs and over 25% training throughput improvement for an ads model on MTIA, alongside a 100% pass rate on all 250 KernelBench problems.

productiongpu-kernelsmtiatritonagents