KernelEvolve
KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at MetaMeta's agentic system that takes kernel specifications and generates and tunes kernels in Triton, CuTe DSL and lower-level languages for recommendation model training and inference on NVIDIA GPUs, AMD GPUs and Meta's MTIA accelerators, using graph-based search, retrieval-augmented prompts and a hardware knowledge base.
LLM agents write and benchmark kernels for Meta's production ranking models, and a job harness feeds measured performance and diagnostics back to the agent over hundreds of candidates. The kernels speed up training and inference of those AI models in production, and Meta says successful strategies are distilled into a reusable skill library and that session data is used to post-train smaller kernel models with RL rewarded by measured kernel performance.
- LLM agents write and benchmark kernels
- Kernels speed up AI model training
- Session data post-trains smaller kernel models
- LLM agents write and benchmark kernels
- Kernels speed up AI model training
- Session data post-trains smaller kernel models
Why it is a road to recursion
It is a production case where agent-written kernels speed up a company's own model training and inference, and the agent's experience is recycled into skills and training data for the next kernel models.
Evidence
Meta reports over 60% inference throughput improvement for its Andromeda ads model on NVIDIA GPUs and over 25% training throughput improvement for an ads model on MTIA, alongside a 100% pass rate on all 250 KernelBench problems.
Related loops
More kernels and compilers →- Gemini models propose code changes
- Evaluators keep best scoring versions
- Finds better Gemini training kernels
- Lowers cost to train next Gemini models
- repeat
- Coding agent edits LLM training code
- Agent benchmarks code on TPU hardware
- Verdict becomes context for next hypothesis
- Wins raise training throughput for same models
- repeat
- Frontier LLMs write CUDA kernels
- Fastest variants seed each new round
- Output feeds next kernel-writing model
- repeat