Alphabell.
Paper · Inference efficiency

Glia

Glia: A Human-Inspired AI for Automated Systems Design and OptimizationGlia is a multi-agent LLM workflow, with reasoning, experimentation and analysis roles, that designs resource-management mechanisms for networked systems by forming hypotheses, running simulations and reading telemetry. Its main case study is a distributed GPU cluster serving LLMs, where it designs request routing, batch scheduling and autoscaling policies.

The loop

LLM agents (OpenAI o3 in the paper's evaluation) write and revise request-routing, scheduling and autoscaling code for an LLM inference cluster, test each design in the Vidur serving simulator, and use the measured completion times and telemetry to form the next hypothesis. The resulting designs lower the latency and GPU cost of serving LLMs, the class of model that runs Glia, and the paper motivates the work by the gap between fast-moving AI models and the infrastructure that serves them.

The loop
Glia
Paper · arXiv
  1. LLM agents write inference cluster code
  2. Designs lower latency and GPU cost
  3. Lowers cost of serving LLMs running Glia
  1. LLM agents write inference cluster code
  2. Designs lower latency and GPU cost
  3. Lowers cost of serving LLMs running Glia
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

Pointed at the cluster that serves its own agents, a cheaper routing or scheduling design buys more agent experiments per dollar for the next design round, which is how serving efficiency compounds into faster AI-driven research.

Evidence

On a request-routing benchmark for a distributed LLM inference system, Glia's router had a mean request completion time 2.2x lower than least-loaded-queue routing and 1.6x lower than OpenEvolve's solution, and Glia took about two hours to match a result a human expert reached in two weeks.

llm-servingrequest-routingschedulingmulti-agentsystems-design