AI makes AI better.
A daily hub for work where AI is part of the loop that improves AI, each entry with its loop spelled out. Read more →
8 picked from 56 candidates · 31 sources read · 56 loops indexed
- 01Loop of the day29 Sep 2026Training data
AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks
The loop: The AREX-2 agent synthesizes long-horizon improvement trajectories from ML engineering tasks and uses them to train its successor for better performance on MLE-bench.
AREX-2 demonstrates that long-horizon reflective data synthesized from ML engineering tasks can effectively train agents to become better at self-improvement.

- 0229 Sep 2026Self-modifying agents
SelfSearch: Reward-Free Search for Self-Improving Agents
The loop: An LLM agent modifies its own instructions and tools using records of its own previous attempts, improving its success rate on future tasks.

- 0317 Apr 2026Self-modifying agents
ModernOps888/the-forge
The loop: Four LLMs compete to evolve and breed code through a JIT compiler judge, creating a closed loop of code improvement that enhances their own capabilities.

- 0428 Sep 2026Self-modifying agents
algorithmicsuperintelligence/openevolve: v0.4.0
The loop: The OpenEvolve system uses pluggable selection strategies and program validation to evolve code, capturing token usage and enforcing evolution blocks for its own mutations.

- 052 Oct 2026Kernels and compilers
Building a High-Performance and Portable vLLM Linear Backend with Helion
The loop: The Helion autotuner uses a DSL to generate high-performance kernels for vLLM, which then runs LLM inference more efficiently for future tasks.

- 0629 Sep 2026Self-reward and self-play
RLTL;DR: Self-improvement by Internalizing Self-generated Feedback
The loop: A policy model generates its own feedback insights from failed attempts and internalizes them through training, improving its success rate on future rollouts.

- 0713 Aug 2026Tools and harness
ZK-Andy/dsh-continual-evolve
The loop: A continual self-evolution plugin refines the DeepSeek Harness state based on session trajectories, using a benchmark-driven loop to validate and improve its own harness.

- 0829 Sep 2026Tools and harness
Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer
The loop: A frozen model acts as a solver to generate run records and then as a proposer to edit its own harness, improving its performance on subsequent tasks.

Loop index
All 56 entries →
- Gemini models propose code changes
- Evaluators keep best scoring versions
- Finds better Gemini training kernels
- Lowers cost to train next Gemini models
- repeat

- Coding agent edits its own code
- Scored child agents enter the archive
- Archived agents make later self modifications
- repeat

- RL agent lays out TPU blocks
- Better layouts make AI hardware cheaper
- Agent pre-trains on earlier chip generations
- repeat

- Coding agent edits LLM training script
- Runs a 5-minute training job
- Keeps change if validation improves
- Changes transferred to larger models
- repeat

- Research agent writes harness mutation blueprints
- Coding agent implements harness module changes
- Winning harness becomes the new champion
- New champion harness runs further generations
- repeat

- Aligned model generates training data
- Reward model judges and filters it
- Data trains next intermediate instruct model
- Next model generates better alignment data
- repeat

- Model generates candidate responses to prompts
- Model scores them itself as a judge
- Scores become preference pairs for DPO
- Trained model becomes next generator and judge
- repeat

- Claude Code writes its own code changes
- Changes become next Claude Code versions
- Release becomes harness for next development round
- repeat

- LLM agents write and benchmark kernels
- Kernels speed up AI model training
- Session data post-trains smaller kernel models
- repeat

- Coding agent writes fine-tuning code
- Agent post-trains a base LLM
- Trained model scored on target benchmark
- One turn of model training another
- repeat

- Learned optimizers update other learned optimizers
- Training keeps best performing optimizer parameters
- Better optimizers train each other faster
- repeat

- LLM reads reasoning model thoughts
- LLM flags reward hacking behavior
- Signal feeds back into training reward
- Trains more aligned capable agents
- repeat
Support for loop research
How to apply →Grants
Funding for projects that put AI in the loop of making AI better. Rolling, low paperwork, sized to the work.
About grants →Fellowships
Funded time, compute and mentorship for researchers who want to go deep on one loop.
About fellowships →Compute
GPU time and credits for loop experiments, which almost always have to run more than once.
About compute →Have a loop worth building?
Tell us what the AI improves and how the improvement comes back. Applications for grants, fellowships and compute all start on the contact page.