AI makes AI better.
A daily hub for work where AI is part of the loop that improves AI, each entry with its loop spelled out. Read more →
6 picked from 127 candidates · 31 sources read · 59 loops indexed
- 01Loop of the day8 Oct 2026Research automation
FreeEvolve: Learning to Evolve Beyond Fixed Loops
The loop: The FREEEVOLVE agent automates the design of workflows and evaluation loops for language model agents. It improves its own evolution skill through meta-evolution by scoring each candidate skill on the fresh target agent it produces.
FreeEvolve takes a significant step by automating not just the agent's prompts or skills, but the entire evolutionary search loop used to discover them.

- 026 Oct 2026Tools and harness
PhysEvo: Astra Can Act, Let It
The loop: A meta-agent diagnoses failures in robot manipulation tasks to revise its tools and skills. It also improves its own diagnostic tools, ensuring that retained revisions support both later action and subsequent self-improvement.

- 037 Oct 2026AI R&D evals
EBR-bench update
The loop: EBR-bench measures the ability of models to learn from their own experience. By evaluating how well models adapt, the benchmark provides a metric that tracks and guides the development of self-improving AI systems.
loop fit 8/10via Epoch AI
- 0424 Jul 2026Self-modifying agents
Birfy/agentdescent
The loop: The agentdescent framework treats agent parameters like skills and harnesses as optimizable components. This allows an optimizer to compute diffs as gradients and update the agents, enabling them to evolve their own components to improve performance.

- 057 Oct 2026Training data
FrogNano: Training a 4B Coding Agent via Online Task Synthesis
The loop: An online task synthesis pipeline creates coding tasks calibrated to a 4B agent's current learnability frontier. The agent trains on these synthetic tasks via reinforcement learning, which shifts its frontier and prompts the pipeline to generate harder tasks for the next round.

- 06Tools and harness
AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation
The loop: AutoRef automates the optimization of evaluation harnesses for image generation agents, improving the accuracy and efficiency of model development.
loop fit 9/10via Awesome-AI4AI
- 076 Oct 2026Tools and harness
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
The loop: An LLM agent system iteratively proposes and selects edits to its own harness using a regularized proposer and critic. The resulting improved harness directly upgrades the agent's operating environment, making it better at solving tasks and further improving its harness.

- 085 Oct 2026Tools and harness
Second-Order Problem Solving for Recursive Self-Improvement in Formal Verification
The loop: An agent framework monitors its own execution traces to diagnose structural failures and systematically edit its workflow. These edits improve the agent's workflow for formal verification, making it more effective in subsequent rounds of recursive self-improvement.

Loop index
All 59 entries →
- Gemini models propose code changes
- Evaluators keep best scoring versions
- Finds better Gemini training kernels
- Lowers cost to train next Gemini models
- repeat

- Coding agent edits its own code
- Scored child agents enter the archive
- Archived agents make later self modifications
- repeat

- RL agent lays out TPU blocks
- Better layouts make AI hardware cheaper
- Agent pre-trains on earlier chip generations
- repeat

- Coding agent edits LLM training script
- Runs a 5-minute training job
- Keeps change if validation improves
- Changes transferred to larger models
- repeat

- Research agent writes harness mutation blueprints
- Coding agent implements harness module changes
- Winning harness becomes the new champion
- New champion harness runs further generations
- repeat

- Aligned model generates training data
- Reward model judges and filters it
- Data trains next intermediate instruct model
- Next model generates better alignment data
- repeat

- Model generates candidate responses to prompts
- Model scores them itself as a judge
- Scores become preference pairs for DPO
- Trained model becomes next generator and judge
- repeat

- Claude Code writes its own code changes
- Changes become next Claude Code versions
- Release becomes harness for next development round
- repeat

- LLM agents write and benchmark kernels
- Kernels speed up AI model training
- Session data post-trains smaller kernel models
- repeat

- Coding agent writes fine-tuning code
- Agent post-trains a base LLM
- Trained model scored on target benchmark
- One turn of model training another
- repeat

- Learned optimizers update other learned optimizers
- Training keeps best performing optimizer parameters
- Better optimizers train each other faster
- repeat

- LLM reads reasoning model thoughts
- LLM flags reward hacking behavior
- Signal feeds back into training reward
- Trains more aligned capable agents
- repeat
Support for loop research
How to apply →Grants
Funding for projects that put AI in the loop of making AI better. Rolling, low paperwork, sized to the work.
About grants →Fellowships
Funded time, compute and mentorship for researchers who want to go deep on one loop.
About fellowships →Compute
GPU time and credits for loop experiments, which almost always have to run more than once.
About compute →Have a loop worth building?
Tell us what the AI improves and how the improvement comes back. Applications for grants, fellowships and compute all start on the contact page.