Paper · Inference efficiency
SEIS
SEIS: Self-Evolving Inference SystemsAn agentic system autonomously optimizes the mini-sglang inference engine end-to-end through iterative code changes and inherited experiences, achieving a 3.27X throughput speedup.

The loop
SEIS autonomously optimizes its own inference engine code through iterative self-evolution. This redesigns the engine for higher throughput, which accelerates the models that power the system.
The loop
SEIS- Agent edits inference engine code
- Engine is tested for throughput
- Agent inherits experiences for next session
- Faster engine serves the agent
- Agent edits inference engine code
- Engine is tested for throughput
- Agent inherits experiences for next session
- Faster engine serves the agent
↻ The improved system does the next round, and the loop turns again.
Why it is a road to recursion
End-to-end autonomous optimization of inference systems creates a direct feedback loop where an agent makes its own execution faster and cheaper.
Evidence
Serving Qwen3-0.6B on H100, the resulting engine reaches 3.27X the throughput of the original mini-sglang implementation.
inferencecode-generationagent
Related loops
More inference efficiency →
Context Language Models
- Model updates context file
- Skill-optimization evolves instructions
- Model uses improved instructions
- Model manages context better
- repeat
Inference efficiencyRu-Lin Shao et al. · 2026

FlashInfer-Bench
- LLM agents write GPU kernels
- Faster kernels make LLM inference cheaper
- New serving traces define next kernel tasks
- repeat
Inference efficiencyUW, CMU, NVIDIA, UC Berkeley (Xing et al.) · 2026

Glia
- LLM agents write inference cluster code
- Designs lower latency and GPU cost
- Lowers cost of serving LLMs running Glia
- repeat
Inference efficiencyHamadanian et al. (MIT CSAIL) · 2025