Paper · Inference efficiency
Context Language Models
Language models are designed to natively manage their own context as a file, enabling intrinsic context management that can be optimized through skill evolution.
The loop
A language model natively manages its own context by updating a context file. The model's context-management instructions are then evolved through a skill-optimization loop, improving its accuracy and efficiency on subsequent tasks.
The loop
Context Language Models- Model updates context file
- Skill-optimization evolves instructions
- Model uses improved instructions
- Model manages context better
- Model updates context file
- Skill-optimization evolves instructions
- Model uses improved instructions
- Model manages context better
↻ The improved system does the next round, and the loop turns again.
Why it is a road to recursion
Shifting context management from external harnesses to intrinsic model behavior allows models to learn and optimize their own memory strategies.
Evidence
Improves held-out accuracy by up to 35.9 points and achieves 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus.
context-managementskill-evolutionefficiency
Related loops
More inference efficiency →FlashInfer-Bench
- LLM agents write GPU kernels
- Faster kernels make LLM inference cheaper
- New serving traces define next kernel tasks
- repeat
Inference efficiencyUW, CMU, NVIDIA, UC Berkeley (Xing et al.) · 2026
Glia
- LLM agents write inference cluster code
- Designs lower latency and GPU cost
- Lowers cost of serving LLMs running Glia
- repeat
Inference efficiencyHamadanian et al. (MIT CSAIL) · 2025
ADRS
- LLMs propose and mutate load balancer code
- Best programs seed the next generation
- Output is faster serving code for LLMs
- repeat
Inference efficiencyUC Berkeley (Sky Computing Lab) · 2025