ADRS
Barbarians at the Gate: How AI is Upending Systems ResearchThe paper argues that systems research suits AI-driven discovery because performance problems come with reliable verifiers (simulators and workloads), names the approach AI-Driven Research for Systems (ADRS), and runs the open-source OpenEvolve framework on case studies that include Mixture-of-Experts inference load balancing and LLM inference over SQL queries.
LLMs inside OpenEvolve (Gemini 2.5 Flash and Flash Lite for the MoE case) propose and mutate code for the expert-parallelism load balancer that places MoE expert replicas on GPUs during LLM inference, and a 168-line PyTorch simulator scores each candidate on load imbalance and rebalancing runtime, with the best programs seeding the next generation. A second case evolves the row and field reordering that maximizes prefix KV-cache reuse when SQL queries call an LLM on every row. The output is faster serving code for LLMs, the same class of model that wrote it, which is the indirect loop.
- LLMs propose and mutate load balancer code
- Best programs seed the next generation
- Output is faster serving code for LLMs
- LLMs propose and mutate load balancer code
- Best programs seed the next generation
- Output is faster serving code for LLMs
Why it is a road to recursion
Because systems problems come with cheap, reliable verifiers, the same evolve-and-measure loop can be pointed at the serving stack of the LLMs that do the evolving, so each inference-efficiency gain makes the next round of search cheaper.
Evidence
For MoE expert-parallelism load balancing, the evolved algorithm matched the baselines' imbalance factor while cutting rebalancing runtime to 3.7 ms, a 5.0x speedup over a non-public frontier-lab reference implementation (19.6 ms), in a run of about five hours that cost under $10.
Related loops
More inference efficiency →- Model updates context file
- Skill-optimization evolves instructions
- Model uses improved instructions
- Model manages context better
- repeat
- LLM agents write GPU kernels
- Faster kernels make LLM inference cheaper
- New serving traces define next kernel tasks
- repeat
- LLM agents write inference cluster code
- Designs lower latency and GPU cost
- Lowers cost of serving LLMs running Glia
- repeat