Glia
Glia: A Human-Inspired AI for Automated Systems Design and OptimizationGlia is a multi-agent LLM workflow, with reasoning, experimentation and analysis roles, that designs resource-management mechanisms for networked systems by forming hypotheses, running simulations and reading telemetry. Its main case study is a distributed GPU cluster serving LLMs, where it designs request routing, batch scheduling and autoscaling policies.
LLM agents (OpenAI o3 in the paper's evaluation) write and revise request-routing, scheduling and autoscaling code for an LLM inference cluster, test each design in the Vidur serving simulator, and use the measured completion times and telemetry to form the next hypothesis. The resulting designs lower the latency and GPU cost of serving LLMs, the class of model that runs Glia, and the paper motivates the work by the gap between fast-moving AI models and the infrastructure that serves them.
- LLM agents write inference cluster code
- Designs lower latency and GPU cost
- Lowers cost of serving LLMs running Glia
- LLM agents write inference cluster code
- Designs lower latency and GPU cost
- Lowers cost of serving LLMs running Glia
Why it is a road to recursion
Pointed at the cluster that serves its own agents, a cheaper routing or scheduling design buys more agent experiments per dollar for the next design round, which is how serving efficiency compounds into faster AI-driven research.
Evidence
On a request-routing benchmark for a distributed LLM inference system, Glia's router had a mean request completion time 2.2x lower than least-loaded-queue routing and 1.6x lower than OpenEvolve's solution, and Glia took about two hours to match a result a human expert reached in two weeks.
Related loops
More inference efficiency →- Model updates context file
- Skill-optimization evolves instructions
- Model uses improved instructions
- Model manages context better
- repeat
- LLM agents write GPU kernels
- Faster kernels make LLM inference cheaper
- New serving traces define next kernel tasks
- repeat
- LLMs propose and mutate load balancer code
- Best programs seed the next generation
- Output is faster serving code for LLMs
- repeat