The hub
Loop index
Projects, papers, benchmarks, datasets and mini-projects where AI is part of the loop that improves AI. Each entry states its loop and how far it has turned.
Showing 56 entries.

AlphaEvolve
- Gemini models propose code changes
- Evaluators keep best scoring versions
- Finds better Gemini training kernels
- Lowers cost to train next Gemini models
- repeat
Kernels and compilersGoogle DeepMind · 2025

Darwin Gödel Machine
- Coding agent edits its own code
- Scored child agents enter the archive
- Archived agents make later self modifications
- repeat
Self-modifying agentsUBC, Vector Institute, Sakana AI · 2025

AlphaChip
- RL agent lays out TPU blocks
- Better layouts make AI hardware cheaper
- Agent pre-trains on earlier chip generations
- repeat
Hardware and chipsGoogle DeepMind, Google Research · 2020

autoresearch
- Coding agent edits LLM training script
- Runs a 5-minute training job
- Keeps change if validation improves
- Changes transferred to larger models
- repeat
Research automationAndrej Karpathy · 2026

ScholarEvolve
- Research agent writes harness mutation blueprints
- Coding agent implements harness module changes
- Winning harness becomes the new champion
- New champion harness runs further generations
- repeat
Tools and harnessYang et al. (UC Santa Barbara, Microsoft) · 2026

Nemotron-4 340B
- Aligned model generates training data
- Reward model judges and filters it
- Data trains next intermediate instruct model
- Next model generates better alignment data
- repeat
Training dataNVIDIA · 2024

Self-Rewarding LMs
- Model generates candidate responses to prompts
- Model scores them itself as a judge
- Scores become preference pairs for DPO
- Trained model becomes next generator and judge
- repeat
Self-reward and self-playMeta, NYU (Yuan et al.) · 2024

Claude Code
- Claude Code writes its own code changes
- Changes become next Claude Code versions
- Release becomes harness for next development round
- repeat
Self-modifying agentsAnthropic · 2025

KernelEvolve
- LLM agents write and benchmark kernels
- Kernels speed up AI model training
- Session data post-trains smaller kernel models
- repeat
Kernels and compilersMeta · 2025

PostTrainBench
- Coding agent writes fine-tuning code
- Agent post-trains a base LLM
- Trained model scored on target benchmark
- One turn of model training another
- repeat
AI R&D evalsELLIS Institute Tubingen, MPI-IS, University of Tubingen, Thoughtful Lab (Rank et al.) · 2026

Self-training learned optimizers
- Learned optimizers update other learned optimizers
- Training keeps best performing optimizer parameters
- Better optimizers train each other faster
- repeat
Architecture and optimizer searchMetz et al. (Google Research, Brain Team) · 2021

CoT monitoring
- LLM reads reasoning model thoughts
- LLM flags reward hacking behavior
- Signal feeds back into training reward
- Trains more aligned capable agents
- repeat
Interpretability and oversightOpenAI (Baker et al.) · 2025

Context Language ModelsNew
- Model updates context file
- Skill-optimization evolves instructions
- Model uses improved instructions
- Model manages context better
- repeat
Inference efficiencyRu-Lin Shao et al. · 2026

AIDE2
- Outer agent rewrites inner agent code
- Accepted rewrites become the incumbent agent
- Discovered agent becomes the outer loop agent
- repeat
Self-modifying agentsWeco AI · 2026

Redwood
- AI system writes accelerator hardware
- Qwen runs on the new accelerator
- Qwen finds optimizations for the accelerator
- Optimizations improve next accelerator generation
- repeat
Hardware and chipsArchitect Labs · 2026

TPU autoresearch
- Coding agent edits LLM training code
- Agent benchmarks code on TPU hardware
- Verdict becomes context for next hypothesis
- Wins raise training throughput for same models
- repeat
Kernels and compilersAleksey Vlasenko · 2026

Self-optimizing Deep Research
- LLM optimizers critique deep research reports
- Optimizers rewrite research agent prompts
- Optimized system runs the next queries
- repeat
Tools and harnessZeta Alpha · 2026

autoresearch-distillation
- Qwen3-14B edits GPT training script
- Edit is trained and scored
- Score rewards Qwen3-14B weights update
- Trained checkpoint re-run in autoresearch loop
- repeat
Research automationExperiential Labs (Naihin, Fallah) · 2026

FlashInfer-Bench
- LLM agents write GPU kernels
- Faster kernels make LLM inference cheaper
- New serving traces define next kernel tasks
- repeat
Inference efficiencyUW, CMU, NVIDIA, UC Berkeley (Xing et al.) · 2026

Ricursive Intelligence
- AI systems would design chips
- Chips would train more capable AI
- That AI designs the next chips
- repeat
Hardware and chipsRicursive Intelligence (Anna Goldie, Azalia Mirhoseini) · 2025

Glia
- LLM agents write inference cluster code
- Designs lower latency and GPU cost
- Lowers cost of serving LLMs running Glia
- repeat
Inference efficiencyHamadanian et al. (MIT CSAIL) · 2025

ADRS
- LLMs propose and mutate load balancer code
- Best programs seed the next generation
- Output is faster serving code for LLMs
- repeat
Inference efficiencyUC Berkeley (Sky Computing Lab) · 2025

Petri
- Auditor model tests target model
- LLM judge scores target behavior
- Feeds into next Claude model release
- repeat
Interpretability and oversightAnthropic · 2025

DeepScientist
- LLM agents propose hypotheses on AI tasks
- Agents implement and test them
- Findings Memory steers later proposals
- Validated finding directly improves AI system
- repeat
Research automationWeng et al. (Westlake University) · 2025

Tongyi DeepResearch
- Agent synthesizes agentic pretraining data
- Pipeline adjusts training set in real time
- Data feeds back into model training
- repeat
Training dataTongyi Lab, Alibaba · 2025

ShinkaEvolve
- LLM ensemble rewrites candidate programs
- Evaluator scores and selects best programs
- Evolved mixture-of-experts load-balancing loss
- Loss improves LLM training
- repeat
Architecture and optimizer searchSakana AI · 2025

GEPA
- Reflection LLM writes revised module prompts
- Winning prompt candidates are kept
- Optimized prompts return to same system
- New system traces feed next reflection round
- repeat
Tools and harnessAgrawal et al. (UC Berkeley, Stanford, Databricks, MIT) · 2025

LLM Speedrunning Benchmark
- Agent rewrites LLM training script
- Script makes LLM training faster
- Benchmark scores the recovered speedup
- Speeds up agent's own model training
- repeat
AI R&D evalsMeta and collaborators (Zhao et al.) · 2025

Genesys
- Designer agents propose language model architectures
- Verifier agents pre-train and evaluate designs
- Verified results feed evolutionary population
- Language models search for better architectures
- repeat
Architecture and optimizer searchCheng, Clark, Richardson (Allen Institute for AI, Dartmouth) · 2025

SEAL
- Language model generates a self-edit
- Edit is applied as weight update
- RL rewards self-edit by downstream performance
- repeat
Training dataMIT (Zweiger, Pari et al.) · 2025

Stanford AI-generated kernels
- Frontier LLMs write CUDA kernels
- Fastest variants seed each new round
- Output feeds next kernel-writing model
- repeat
Kernels and compilersStanford CRFM (Ouyang, Mirhoseini, Liang) · 2025

Kevin-32B
- QwQ-32B writes CUDA kernels
- Model refines kernels using compiler feedback
- RL rewards model on measured speedup
- repeat
Kernels and compilersCognition, Stanford · 2025

Absolute Zero
- Model proposes code reasoning tasks
- Model solves them for accuracy reward
- Python executor verifies the answers
- Curriculum shifts as solver improves
- repeat
Self-reward and self-playTsinghua, BIGAI, Penn State (Zhao et al.) · 2025

KernelBook + KernelLLM
- KernelLLM writes Triton GPU kernels
- Faster kernels make AI training cheaper
- Verified kernels train next kernel model
- repeat
Kernels and compilersGPU MODE, Meta · 2025

SWE-smith
- o3-mini injects bugs and writes issues
- SWE-agent solves tasks for expert trajectories
- Trajectories fine-tune Qwen 2.5 Coder
- Open coding agent trained on agent data
- repeat
Training dataStanford, Princeton, Alibaba Qwen (Yang et al.) · 2025

SICA
- Best agent changes its own code
- Edited agent is benchmarked and archived
- Archived agent makes the next edit
- repeat
Self-modifying agentsUniversity of Bristol, iGent AI · 2025

The AI Scientist-v2
- LLM agents propose ML research ideas
- Agents run experiments with tree search
- Agents write up results and papers
- Output is new knowledge about models
- repeat
Research automationYamada et al. (Sakana AI) · 2025

PaperBench
- Agent replicates published ML research
- LLM judge grades the replication
- Measures ability to do ML research
- Research produces better future models
- repeat
AI R&D evalsOpenAI · 2025

rStar-Math
- Policy model runs Monte Carlo Tree Search
- Verified trajectories train next policy and PPM
- PPM guides search to produce better data
- Self-evolution improves both models
- repeat
Self-reward and self-playMicrosoft Research Asia · 2025

KernelBench
- LLM writes replacement GPU kernel
- Kernel is timed against baseline
- Workloads are neural network operators
- Makes agent's own training faster
- repeat
AI R&D evalsOuyang et al. (Stanford) · 2024

RE-Bench
- Agents do AI research tasks
- Scored against human expert baselines
- Measures automation of AI research
- Research improves future AI systems
- repeat
AI R&D evalsMETR · 2024

MLE-bench
- Agents build and train ML models
- Submissions scored against human thresholds
- Measures open-ended ML research ability
- Agents improve their own training code
- repeat
AI R&D evalsOpenAI · 2024

ADAS (Meta Agent Search)
- Meta agent programs new agent designs
- Designs are evaluated on target tasks
- Results are added to an archive
- Meta agent uses archive for next design
- repeat
Tools and harnessHu, Lu, Clune (UBC, Vector Institute) · 2024

Self-Taught Evaluator
- LLM builds worse answers for preference pairs
- Judge samples reasoning traces and verdicts
- Judge is fine-tuned on correct verdicts
- Improved judge labels next iteration's data
- repeat
Self-reward and self-playMeta FAIR · 2024

CriticGPT
- CriticGPT critiques ChatGPT code answers
- Humans catch errors using critiques
- Improves labels for RLHF training
- Labels train the next ChatGPT
- repeat
Interpretability and oversightOpenAI (McAleese et al.) · 2024

DiscoPOP
- GPT-4 proposes preference-optimization losses
- Candidate losses fine-tune language models
- Evaluation scores feed back into prompt
- Best loss aligns other LLMs
- repeat
Architecture and optimizer searchSakana AI, FLAIR (Oxford), University of Cambridge · 2024

FineWeb-Edu
- Llama-3-70B scores pages for educational value
- Regressor trained on labels filters corpus
- Filtered corpus pretrains new language models
- LLM judgment sets next training diet
- repeat
Training dataHugging Face · 2024

MAIA
- Agent runs interpretability experiments
- Agent finds spurious background cues
- Selection retrains classifier final layer
- Improves robustness of interpreted model
- repeat
Interpretability and oversightMIT CSAIL (Rott Shaham et al.) · 2024

STOP
- GPT-4 rewrites the improver program
- Each rewrite is scored by meta-utility
- Best rewrite becomes next round improver
- repeat
Self-modifying agentsZelikman, Lorch, Mackey, Kalai (Stanford, Microsoft Research, OpenAI) · 2023

Aider
- Aider writes new code for Aider releases
- New code becomes the next release
- Next release builds the following version
- repeat
Self-modifying agentsPaul Gauthier (Aider) · 2023

Model-written evals
- Language model writes evaluation questions
- Evaluations test other language models
- Exposes behaviors worsened by RLHF
- Evaluates next round of training
- repeat
Interpretability and oversightAnthropic (Perez et al.) · 2022

Constitutional AI
- Model critiques and revises its own responses
- Model is fine-tuned on the revisions
- Model judges responses to train preference model
- Preference model provides RL reward signal
- repeat
Self-reward and self-playBai et al. (Anthropic) · 2022

STaR
- Language model writes step-by-step rationales
- Correct rationales are kept as dataset
- Model is fine-tuned on self-generated set
- Improved model generates next rationale dataset
- repeat
Training dataZelikman et al. (Stanford, Google Research) · 2022

AlphaZero
- Network plays games against itself
- Games become the training data
- Same network trains on its own moves
- Updated network plays next self-play games
- repeat
Self-reward and self-playSilver et al. (DeepMind) · 2017

NAS with RL
- RNN controller writes child network architecture
- Child validation accuracy returns as reward
- Updates make controller propose better architectures
- repeat
Architecture and optimizer searchZoph and Le (Google Brain) · 2016

Gödel Machine
- Proof searcher tests proof techniques
- Executes useful rewrites of its own code
- Rewritten machine searches on its new code
- repeat
Self-modifying agentsSchmidhuber (IDSIA) · 2003
Nothing matches. Try another filter, or submit a project.
Missing a loop?
Send us a project, paper or repo with the loop in one sentence and we will review it for the index.