AI makes AI better.
A daily hub for work where AI is part of the loop that improves AI, each entry with its loop spelled out. Read more →
7 picked from 5 candidates · 31 sources read · 56 loops indexed
- 01Loop of the day27 Sep 2026Self-modifying agents
Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents
The loop: The Rep2Skill agent models its own internal representation trajectories to localize execution errors, generating textual feedback to evolve its external skills.
Rep2Skill pushes skill evolution beyond text-only feedback by allowing agents to reflect directly on their internal representation trajectories to diagnose and fix execution errors.

- 0229 Sep 2026Inference efficiency
AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents
The loop: A coding agent is trained to decide when and how to compact its own context during long-horizon tasks, using task-success rewards to improve its policy.

- 0328 Sep 2026Hardware and chips
How NVIDIA DSX MaxLPS Maximizes AI Factory Throughput and Efficiency
The loop: DSX MaxLPS optimizes power usage in AI factories, which improves throughput for the models running in that factory.

- 041 Oct 2026Tools and harness
LiteEvo: Automated, Cost-Efficient Harness Evolution for Generalization to Unseen Tasks
The loop: LiteEvo meta-agents mine agent trajectories for reusable components to evolve a harness library, which improves the agent's performance on unseen tasks.

- 0530 Sep 2026Tools and harness
From Upstream Changes to Downstream Confidence: Inside Torch Spyre’s Integration with PyTorch CRCR
The loop: Torch Spyre uses the CRCR relay to automate CI testing, which improves the stability of PyTorch for its own development.
loop fit 7/10Mehant Kammakomati (IBM), Jewel K M (Red Hat), Anubhav Jana (IBM), Padmanabha...via PyTorch
- 0630 Sep 2026Self-reward and self-play
AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation
The loop: An advisor model uses reflection to propose corrections to its own decisions, then uses self-distillation from a feedback-conditioned copy to improve its future advice.

- 0728 Sep 2026Interpretability and oversight
Towards safety cases for frontier AI training
The loop: A safety case framework investigates misalignment incidents, which improves technical safeguards for the training of successor frontier models.
loop fit 6/10via OpenAI
- 0829 Sep 2026Tools and harness
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
The loop: MILO mutator agents rewrite complete agent harnesses and receive parent-specific feedback on their performance. An orchestrator uses this global search history to adapt the mutators' assignments and curriculum, improving the discovery of future harnesses.

Loop index
All 56 entries →
- Gemini models propose code changes
- Evaluators keep best scoring versions
- Finds better Gemini training kernels
- Lowers cost to train next Gemini models
- repeat

- Coding agent edits its own code
- Scored child agents enter the archive
- Archived agents make later self modifications
- repeat

- RL agent lays out TPU blocks
- Better layouts make AI hardware cheaper
- Agent pre-trains on earlier chip generations
- repeat

- Coding agent edits LLM training script
- Runs a 5-minute training job
- Keeps change if validation improves
- Changes transferred to larger models
- repeat

- Research agent writes harness mutation blueprints
- Coding agent implements harness module changes
- Winning harness becomes the new champion
- New champion harness runs further generations
- repeat

- Aligned model generates training data
- Reward model judges and filters it
- Data trains next intermediate instruct model
- Next model generates better alignment data
- repeat

- Model generates candidate responses to prompts
- Model scores them itself as a judge
- Scores become preference pairs for DPO
- Trained model becomes next generator and judge
- repeat

- Claude Code writes its own code changes
- Changes become next Claude Code versions
- Release becomes harness for next development round
- repeat

- LLM agents write and benchmark kernels
- Kernels speed up AI model training
- Session data post-trains smaller kernel models
- repeat

- Coding agent writes fine-tuning code
- Agent post-trains a base LLM
- Trained model scored on target benchmark
- One turn of model training another
- repeat

- Learned optimizers update other learned optimizers
- Training keeps best performing optimizer parameters
- Better optimizers train each other faster
- repeat

- LLM reads reasoning model thoughts
- LLM flags reward hacking behavior
- Signal feeds back into training reward
- Trains more aligned capable agents
- repeat
Support for loop research
How to apply →Grants
Funding for projects that put AI in the loop of making AI better. Rolling, low paperwork, sized to the work.
About grants →Fellowships
Funded time, compute and mentorship for researchers who want to go deep on one loop.
About fellowships →Compute
GPU time and credits for loop experiments, which almost always have to run more than once.
About compute →Have a loop worth building?
Tell us what the AI improves and how the improvement comes back. Applications for grants, fellowships and compute all start on the contact page.