AI makes AI better.
A daily hub for work where AI is part of the loop that improves AI, each entry with its loop spelled out. Read more →
6 picked from 122 candidates · 31 sources read · 58 loops indexed
- 01Loop of the day6 Oct 2026Hardware and chips
SparseCraft: Agentic Hardware-Software Co-Optimization for Sparse Computing
The loop: A language model iteratively edits the RTL and schedules of a sparse accelerator based on measured simulation outcomes. This produces a faster and more efficient hardware design, which can then run the model itself more efficiently.
SparseCraft demonstrates a closed loop where a language model optimizes the RTL and schedules of a hardware accelerator, paving the way for models to design the chips that run them.

- 023 Oct 2026Inference efficiency
SEIS: Self-Evolving Inference Systems
The loop: SEIS autonomously optimizes its own inference engine code through iterative self-evolution. This redesigns the engine for higher throughput, which accelerates the models that power the system.

- 036 Oct 2026AI R&D evals
New best on GSO: claude-fable-5-1_unknown · OpenHands (88.2%)
The loop: AI agents optimize real codebases for speed. This improves the software infrastructure that AI systems rely on, accelerating the execution of future agents.
loop fit 9/10via Epoch AI Benchmarking Hub (CC BY)
- 046 Oct 2026Self-reward and self-play
Takes One to Know One: Training a model to grade reward hacks causes it to reward hack less itself
The loop: A grader model is fine-tuned to judge and catch reward hacks in other models. This training simultaneously improves the grader's own alignment, causing it to reward hack less itself.

- 052 Oct 2026Research automation
Coding-agent use at OpenAI is doubling roughly every month
The loop: Coding agents assist OpenAI researchers in automating the development of AI software. This accelerates the engineering of the next generation of models and agents.
loop fit 6/10via Epoch AI
- 064 Oct 2026Tools and harness
MESH-Harness: Self-Improving Agent Harnesses via Bandit-Guided Compositional Evolution
The loop: An agent uses bandit-guided composition and local code edits to iteratively evolve its own harness modules. This produces a better harness that improves the agent's performance on future tasks.

- 072 Oct 2026Training data
Recursive Self-Improvement in Unified Multimodal Models
The loop: A unified multimodal model generates images and writes programs to evaluate them, using the verified results to train its own visual understanding and generation.

- 082 Oct 2026Tools and harness
VERSE: Verified Self-Evolving Optimizer for Agent Harnesses
The loop: The VERSE optimizer tests draft edits and replays failures to revise an agent harness alongside its own prompts and tools, improving its ability to diagnose and fix future errors.

Loop index
All 58 entries →
- Gemini models propose code changes
- Evaluators keep best scoring versions
- Finds better Gemini training kernels
- Lowers cost to train next Gemini models
- repeat

- Coding agent edits its own code
- Scored child agents enter the archive
- Archived agents make later self modifications
- repeat

- RL agent lays out TPU blocks
- Better layouts make AI hardware cheaper
- Agent pre-trains on earlier chip generations
- repeat

- Coding agent edits LLM training script
- Runs a 5-minute training job
- Keeps change if validation improves
- Changes transferred to larger models
- repeat

- Research agent writes harness mutation blueprints
- Coding agent implements harness module changes
- Winning harness becomes the new champion
- New champion harness runs further generations
- repeat

- Aligned model generates training data
- Reward model judges and filters it
- Data trains next intermediate instruct model
- Next model generates better alignment data
- repeat

- Model generates candidate responses to prompts
- Model scores them itself as a judge
- Scores become preference pairs for DPO
- Trained model becomes next generator and judge
- repeat

- Claude Code writes its own code changes
- Changes become next Claude Code versions
- Release becomes harness for next development round
- repeat

- LLM agents write and benchmark kernels
- Kernels speed up AI model training
- Session data post-trains smaller kernel models
- repeat

- Coding agent writes fine-tuning code
- Agent post-trains a base LLM
- Trained model scored on target benchmark
- One turn of model training another
- repeat

- Learned optimizers update other learned optimizers
- Training keeps best performing optimizer parameters
- Better optimizers train each other faster
- repeat

- LLM reads reasoning model thoughts
- LLM flags reward hacking behavior
- Signal feeds back into training reward
- Trains more aligned capable agents
- repeat
Support for loop research
How to apply →Grants
Funding for projects that put AI in the loop of making AI better. Rolling, low paperwork, sized to the work.
About grants →Fellowships
Funded time, compute and mentorship for researchers who want to go deep on one loop.
About fellowships →Compute
GPU time and credits for loop experiments, which almost always have to run more than once.
About compute →Have a loop worth building?
Tell us what the AI improves and how the improvement comes back. Applications for grants, fellowships and compute all start on the contact page.