AI makes AI better.
A daily hub for work where AI is part of the loop that improves AI, each entry with its loop spelled out. Read more →
5 picked from 128 candidates · 31 sources read · 58 loops indexed
- 01Loop of the day6 Oct 2026Tools and harness
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
The loop: An LLM agent system iteratively proposes and selects edits to its own harness using a regularized proposer and critic. The resulting improved harness directly upgrades the agent's operating environment, making it better at solving tasks and further improving its harness.
RRSI demonstrates how regularizing the self-improvement process can prevent an agent from overfitting to its training tasks while evolving its own harness.

- 025 Oct 2026Tools and harness
Second-Order Problem Solving for Recursive Self-Improvement in Formal Verification
The loop: An agent framework monitors its own execution traces to diagnose structural failures and systematically edit its workflow. These edits improve the agent's workflow for formal verification, making it more effective in subsequent rounds of recursive self-improvement.

- 036 Oct 2026AI R&D evals
New best on GSO: claude-fable-5-1_unknown · OpenHands (88.2%)
The loop: AI agents optimize real codebases for speed. This improves the software infrastructure that AI systems rely on, accelerating the execution of future agents.
loop fit 9/10via Epoch AI Benchmarking Hub (CC BY)
- 046 Oct 2026Self-reward and self-play
Takes One to Know One: Training a model to grade reward hacks causes it to reward hack less itself
The loop: A grader model is fine-tuned to judge and catch reward hacks in other models. This training simultaneously improves the grader's own alignment, causing it to reward hack less itself.

- 053 Oct 2026Architecture and optimizer search
EvoCast: Reliable Autonomous Research Agents for Iterative Forecasting Architecture Evolution
The loop: An autonomous research agent iteratively designs and evaluates time-series forecasting architectures. The experimental outcomes from these evaluations are accumulated as evidence to guide and improve the agent's subsequent architecture search rounds.

- 065 Oct 2026Training data
Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite
The loop: A base model discovers successful solutions under diverse harnesses and rewrites them into training trajectories. These trajectories are then used for supervised finetuning, which directly improves the base model's performance on complex tasks.

- 073 Oct 2026Training data
Trinity: Self-Evolving Vision-Language Models with a Self-Verifier
The loop: A vision-language model acts as a Questioner, Solver, and Verifier to generate and screen its own training data. This verified data is then used to train the model's successor, improving its reasoning capabilities without external labels.

- 083 Oct 2026Inference efficiency
SEIS: Self-Evolving Inference Systems
The loop: SEIS autonomously optimizes its own inference engine code through iterative self-evolution. This redesigns the engine for higher throughput, which accelerates the models that power the system.

Loop index
All 58 entries →
- Gemini models propose code changes
- Evaluators keep best scoring versions
- Finds better Gemini training kernels
- Lowers cost to train next Gemini models
- repeat

- Coding agent edits its own code
- Scored child agents enter the archive
- Archived agents make later self modifications
- repeat

- RL agent lays out TPU blocks
- Better layouts make AI hardware cheaper
- Agent pre-trains on earlier chip generations
- repeat

- Coding agent edits LLM training script
- Runs a 5-minute training job
- Keeps change if validation improves
- Changes transferred to larger models
- repeat

- Research agent writes harness mutation blueprints
- Coding agent implements harness module changes
- Winning harness becomes the new champion
- New champion harness runs further generations
- repeat

- Aligned model generates training data
- Reward model judges and filters it
- Data trains next intermediate instruct model
- Next model generates better alignment data
- repeat

- Model generates candidate responses to prompts
- Model scores them itself as a judge
- Scores become preference pairs for DPO
- Trained model becomes next generator and judge
- repeat

- Claude Code writes its own code changes
- Changes become next Claude Code versions
- Release becomes harness for next development round
- repeat

- LLM agents write and benchmark kernels
- Kernels speed up AI model training
- Session data post-trains smaller kernel models
- repeat

- Coding agent writes fine-tuning code
- Agent post-trains a base LLM
- Trained model scored on target benchmark
- One turn of model training another
- repeat

- Learned optimizers update other learned optimizers
- Training keeps best performing optimizer parameters
- Better optimizers train each other faster
- repeat

- LLM reads reasoning model thoughts
- LLM flags reward hacking behavior
- Signal feeds back into training reward
- Trains more aligned capable agents
- repeat
Support for loop research
How to apply →Grants
Funding for projects that put AI in the loop of making AI better. Rolling, low paperwork, sized to the work.
About grants →Fellowships
Funded time, compute and mentorship for researchers who want to go deep on one loop.
About fellowships →Compute
GPU time and credits for loop experiments, which almost always have to run more than once.
About compute →Have a loop worth building?
Tell us what the AI improves and how the improvement comes back. Applications for grants, fellowships and compute all start on the contact page.