AI makes AI better.
A daily hub for work where AI is part of the loop that improves AI, each entry with its loop spelled out. Read more →
8 picked from 3 candidates · 31 sources read · 56 loops indexed
- 01Loop of the day29 Sep 2026Tools and harness
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
The loop: MILO mutator agents rewrite complete agent harnesses and receive parent-specific feedback on their performance. An orchestrator uses this global search history to adapt the mutators' assignments and curriculum, improving the discovery of future harnesses.
MILO shows how co-evolving an agent's harness alongside the strategy used to discover it can break out of fixed search patterns and yield superior execution environments.

- 0229 Sep 2026Self-modifying agents
Topological Coherence for Self-evolving Multi-agent Systems
The loop: TOCOMAS proposes coupled changes to its own agent, collaboration, and memory policies during online execution. It retains candidates that satisfy structural constraints and improve evaluated reward, directly evolving its own architecture for future tasks.

- 033 Mar 2026Tools and harness
IgorGanapolsky/ThumbGate
The loop: ThumbGate Pre-Action Checks analyze ranked lessons and repeated failures from past executions. The system uses this data to self-improve its strict mode blocking rules, becoming better at preventing secret leaks in future actions.

- 042 Oct 2026Training data
AutoSynthData: Generating Training Data for Enterprise Agents
The loop: The AutoSynthData generator creates synthetic training datasets tailored for enterprise environments. This data is then fed back into the training pipeline to improve the performance and reliability of the enterprise agents.
loop fit 6/10via Hugging Face
- 0529 Sep 2026Self-reward and self-play
Train Ahead, Distill Back: Bootstrapping On-Policy Self-Distillation for Large Language Models
The loop: The B-OPSD policy temporarily trains ahead to create a stronger future teacher. This teacher then generates reliable trajectories and provides dense supervision to the restarted original student, improving the model's own successor.

- 0617 Aug 2026Tools and harness
proteus-evolve/Proteus
The loop: Proteus plugs into any agent harness to measure its performance and propose evolutionary changes. These changes are applied to the harness, improving the execution environment for subsequent agent runs.

- 0726 Sep 2026Training data
X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization
The loop: The X-Tree tokenizer mines flat action streams to build a hierarchy of reusable skills without LLM calls. This tree is then used as a self-teacher during on-policy self-distillation, improving the agent's ability to generalize across tasks.

- 0813 Aug 2026Tools and harness
ruvnet/dream-machine
The loop: The Dream Machine engine schedules nightly repository evolution tasks and evaluates the results. It uses an evidence-gated promotion system to merge successful changes, continuously improving its own configuration and codebase.

Loop index
All 56 entries →
- Gemini models propose code changes
- Evaluators keep best scoring versions
- Finds better Gemini training kernels
- Lowers cost to train next Gemini models
- repeat

- Coding agent edits its own code
- Scored child agents enter the archive
- Archived agents make later self modifications
- repeat

- RL agent lays out TPU blocks
- Better layouts make AI hardware cheaper
- Agent pre-trains on earlier chip generations
- repeat

- Coding agent edits LLM training script
- Runs a 5-minute training job
- Keeps change if validation improves
- Changes transferred to larger models
- repeat

- Research agent writes harness mutation blueprints
- Coding agent implements harness module changes
- Winning harness becomes the new champion
- New champion harness runs further generations
- repeat

- Aligned model generates training data
- Reward model judges and filters it
- Data trains next intermediate instruct model
- Next model generates better alignment data
- repeat

- Model generates candidate responses to prompts
- Model scores them itself as a judge
- Scores become preference pairs for DPO
- Trained model becomes next generator and judge
- repeat

- Claude Code writes its own code changes
- Changes become next Claude Code versions
- Release becomes harness for next development round
- repeat

- LLM agents write and benchmark kernels
- Kernels speed up AI model training
- Session data post-trains smaller kernel models
- repeat

- Coding agent writes fine-tuning code
- Agent post-trains a base LLM
- Trained model scored on target benchmark
- One turn of model training another
- repeat

- Learned optimizers update other learned optimizers
- Training keeps best performing optimizer parameters
- Better optimizers train each other faster
- repeat

- LLM reads reasoning model thoughts
- LLM flags reward hacking behavior
- Signal feeds back into training reward
- Trains more aligned capable agents
- repeat
Support for loop research
How to apply →Grants
Funding for projects that put AI in the loop of making AI better. Rolling, low paperwork, sized to the work.
About grants →Fellowships
Funded time, compute and mentorship for researchers who want to go deep on one loop.
About fellowships →Compute
GPU time and credits for loop experiments, which almost always have to run more than once.
About compute →Have a loop worth building?
Tell us what the AI improves and how the improvement comes back. Applications for grants, fellowships and compute all start on the contact page.