AI makes AI better.
A daily hub for work where AI is part of the loop that improves AI, each entry with its loop spelled out. Read more →
5 picked from 8 candidates · 31 sources read · 61 loops indexed
- 01Loop of the day7 Oct 2026Training data
AdvSim2Real: Training Web Agents Against Adaptive Prompt Injection in a Web World Model
The loop: A web world model co-evolves a task curriculum and an injection adversary to train a web agent. The resulting training data improves the agent's robustness and capability, allowing it to handle more difficult tasks and stronger adversaries in the next iteration.
Co-evolving a task curriculum and an adversary inside a simulated web environment creates a robust training loop that hardens agents against adaptive prompt injections.

- 026 Oct 2026Architecture and optimizer search
RLDiscover: LLM-Driven Co-Evolution of Reinforcement Learning Algorithms
The loop: The RLDiscover framework progressively co-evolves the components of model-free deep reinforcement learning algorithms. The discovered algorithms improve the learning efficiency of the agents that use them, providing better fitness signals for the framework to discover even stronger algorithms.

- 038 Oct 2026Tools and harness
Harness Evolution Hits a Ceiling: When Weight Training Should Begin
The loop: A self-evolving harness loop repairs process failures to generate successful execution trajectories for an LLM agent. These trajectories are then used to train the model's weights, internalizing the gains and making the model better under the original harness.

- 046 Oct 2026Self-modifying agents
Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents
The loop: A self-improving agent constructs meta-experience by re-executing incumbent and revised meta-skills from the same restored discovery state. This hindsight distillation improves the agent's meta-skills, making it better at discovering and refining future task-skills.

- 058 Oct 2026Training data
SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning
The loop: A Synthesizer agent constructs training tasks based on a Reasoner agent's current capabilities, and the Reasoner learns from the resulting experience. The outcomes of the Reasoner's rollouts provide complementary rewards that jointly optimize both agents, making the Synthesizer better at generating informative tasks.

- 069 Oct 2026Self-reward and self-play
The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
The loop: A coder agent and a reviewer agent co-evolve under non-stationary utilities. The reviewer grades patches to guide the coder's search, while the coder's output helps the reviewer refine its grading rubric, making both better at their respective tasks.
loop fit 10/10Alex Iacob, Andrej Jovanovi\'c, William F. Shen, Daniel Burkhardt, Meghdad...via arXiv
- 077 Oct 2026Kernels and compilers
KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization
The loop: A multi-agent system profiles and optimizes generated Triton sub-kernels within compiled models. This improves the execution efficiency of the models and allows the agents to verify the re-stitched model end-to-end.

- 088 Oct 2026Research automation
Language Models as AI Research World Models
The loop: Language models act as Research World Models to predict the outcomes of candidate AI research experiments. This improves the selection of future interventions and generates new experimental data to further train the world model.

Loop index
All 61 entries →
- Gemini models propose code changes
- Evaluators keep best scoring versions
- Finds better Gemini training kernels
- Lowers cost to train next Gemini models
- repeat

- Coding agent edits its own code
- Scored child agents enter the archive
- Archived agents make later self modifications
- repeat

- RL agent lays out TPU blocks
- Better layouts make AI hardware cheaper
- Agent pre-trains on earlier chip generations
- repeat

- Coding agent edits LLM training script
- Runs a 5-minute training job
- Keeps change if validation improves
- Changes transferred to larger models
- repeat

- Research agent writes harness mutation blueprints
- Coding agent implements harness module changes
- Winning harness becomes the new champion
- New champion harness runs further generations
- repeat

- Aligned model generates training data
- Reward model judges and filters it
- Data trains next intermediate instruct model
- Next model generates better alignment data
- repeat

- Model generates candidate responses to prompts
- Model scores them itself as a judge
- Scores become preference pairs for DPO
- Trained model becomes next generator and judge
- repeat

- Claude Code writes its own code changes
- Changes become next Claude Code versions
- Release becomes harness for next development round
- repeat

- LLM agents write and benchmark kernels
- Kernels speed up AI model training
- Session data post-trains smaller kernel models
- repeat

- Coding agent writes fine-tuning code
- Agent post-trains a base LLM
- Trained model scored on target benchmark
- One turn of model training another
- repeat

- Learned optimizers update other learned optimizers
- Training keeps best performing optimizer parameters
- Better optimizers train each other faster
- repeat

- LLM reads reasoning model thoughts
- LLM flags reward hacking behavior
- Signal feeds back into training reward
- Trains more aligned capable agents
- repeat
Support for loop research
How to apply →Grants
Funding for projects that put AI in the loop of making AI better. Rolling, low paperwork, sized to the work.
About grants →Fellowships
Funded time, compute and mentorship for researchers who want to go deep on one loop.
About fellowships →Compute
GPU time and credits for loop experiments, which almost always have to run more than once.
About compute →Have a loop worth building?
Tell us what the AI improves and how the improvement comes back. Applications for grants, fellowships and compute all start on the contact page.