AI makes AI better.
A daily hub for work where AI is part of the loop that improves AI, each entry with its loop spelled out. Read more →
7 picked from 104 candidates · 31 sources read · 57 loops indexed
- 01Loop of the day2 Oct 2026Training data
Recursive Self-Improvement in Unified Multimodal Models
The loop: A unified multimodal model generates images and writes programs to evaluate them, using the verified results to train its own visual understanding and generation.
Cross-capability self-improvement allows a multimodal model to use one modality to verify and generate training data for another, preventing error accumulation.

- 022 Oct 2026Tools and harness
VERSE: Verified Self-Evolving Optimizer for Agent Harnesses
The loop: The VERSE optimizer tests draft edits and replays failures to revise an agent harness alongside its own prompts and tools, improving its ability to diagnose and fix future errors.

- 034 Oct 2026Training data
Reducing Synthetic Markers Makes Some SDF False Facts Linearly Indistinguishable from Pretraining-Acquired Knowledge
The loop: A synthetic document generator produces training data with reduced markers, which finetunes a model to internalize false facts that evade middle-layer linear probes.

- 041 Oct 2026Kernels and compilers
Optimizing Jagged Flash Attention with TLX: The Road Toward SOTA FA4 on Blackwell
The loop: The Jagged Flash Attention kernel improves inference efficiency on Blackwell, which makes the Generative Ads Model faster.
loop fit 6/10Han Xu, Jacky Zhou, Jackie (Jiaqi) Xu, Hongtao Yu, Peng Chen (Dev Infra),...via PyTorch
- 052 Oct 2026Training data
Recursive Harness Self-Improvement for Frontier Reasoning Data Synthesis
The loop: A data synthesis system recursively improves its own generation harness by converting intermediate solver failures into reusable skills, producing progressively harder reasoning tasks.

- 061 Oct 2026Tools and harness
Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
The loop: Olmo-core 3 provides an open training stack that improves the efficiency of training MoE models, which are then used to build better versions of the stack.
loop fit 6/10via Ai2
- 0729 Sep 2026Self-reward and self-play
AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation
The loop: An advisor model reflects on completed interactions to propose corrections, then uses targeted self-distillation to improve the advice it issues to a frozen language-model executor.

- 0827 Sep 2026Self-modifying agents
Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents
The loop: The Rep2Skill agent models its own internal representation trajectories to localize execution errors, generating textual feedback to evolve its external skills.

Loop index
All 57 entries →
- Gemini models propose code changes
- Evaluators keep best scoring versions
- Finds better Gemini training kernels
- Lowers cost to train next Gemini models
- repeat

- Coding agent edits its own code
- Scored child agents enter the archive
- Archived agents make later self modifications
- repeat

- RL agent lays out TPU blocks
- Better layouts make AI hardware cheaper
- Agent pre-trains on earlier chip generations
- repeat

- Coding agent edits LLM training script
- Runs a 5-minute training job
- Keeps change if validation improves
- Changes transferred to larger models
- repeat

- Research agent writes harness mutation blueprints
- Coding agent implements harness module changes
- Winning harness becomes the new champion
- New champion harness runs further generations
- repeat

- Aligned model generates training data
- Reward model judges and filters it
- Data trains next intermediate instruct model
- Next model generates better alignment data
- repeat

- Model generates candidate responses to prompts
- Model scores them itself as a judge
- Scores become preference pairs for DPO
- Trained model becomes next generator and judge
- repeat

- Claude Code writes its own code changes
- Changes become next Claude Code versions
- Release becomes harness for next development round
- repeat

- LLM agents write and benchmark kernels
- Kernels speed up AI model training
- Session data post-trains smaller kernel models
- repeat

- Coding agent writes fine-tuning code
- Agent post-trains a base LLM
- Trained model scored on target benchmark
- One turn of model training another
- repeat

- Learned optimizers update other learned optimizers
- Training keeps best performing optimizer parameters
- Better optimizers train each other faster
- repeat

- LLM reads reasoning model thoughts
- LLM flags reward hacking behavior
- Signal feeds back into training reward
- Trains more aligned capable agents
- repeat
Support for loop research
How to apply →Grants
Funding for projects that put AI in the loop of making AI better. Rolling, low paperwork, sized to the work.
About grants →Fellowships
Funded time, compute and mentorship for researchers who want to go deep on one loop.
About fellowships →Compute
GPU time and credits for loop experiments, which almost always have to run more than once.
About compute →Have a loop worth building?
Tell us what the AI improves and how the improvement comes back. Applications for grants, fellowships and compute all start on the contact page.