Alphabell.
The hub

Loop index

Projects, papers, benchmarks, datasets and mini-projects where AI is part of the loop that improves AI. Each entry states its loop and how far it has turned.

Showing 56 entries.
AlphaEvolve
ProjectGoogle DeepMind
AlphaEvolve
  1. Gemini models propose code changes
  2. Evaluators keep best scoring versions
  3. Finds better Gemini training kernels
  4. Lowers cost to train next Gemini models
  5. repeat
Kernels and compilersGoogle DeepMind · 2025
Darwin Gödel Machine
PaperarXiv
Darwin Gödel Machine
  1. Coding agent edits its own code
  2. Scored child agents enter the archive
  3. Archived agents make later self modifications
  4. repeat
Self-modifying agentsUBC, Vector Institute, Sakana AI · 2025
AlphaChip
ProjectGoogle DeepMind
AlphaChip
  1. RL agent lays out TPU blocks
  2. Better layouts make AI hardware cheaper
  3. Agent pre-trains on earlier chip generations
  4. repeat
Hardware and chipsGoogle DeepMind, Google Research · 2020
autoresearch
Mini-projectGitHub
autoresearch
  1. Coding agent edits LLM training script
  2. Runs a 5-minute training job
  3. Keeps change if validation improves
  4. Changes transferred to larger models
  5. repeat
Research automationAndrej Karpathy · 2026
ScholarEvolve
PaperarXiv
ScholarEvolve
  1. Research agent writes harness mutation blueprints
  2. Coding agent implements harness module changes
  3. Winning harness becomes the new champion
  4. New champion harness runs further generations
  5. repeat
Tools and harnessYang et al. (UC Santa Barbara, Microsoft) · 2026
Nemotron-4 340B
PaperarXiv
Nemotron-4 340B
  1. Aligned model generates training data
  2. Reward model judges and filters it
  3. Data trains next intermediate instruct model
  4. Next model generates better alignment data
  5. repeat
Training dataNVIDIA · 2024
Self-Rewarding LMs
PaperarXiv
Self-Rewarding LMs
  1. Model generates candidate responses to prompts
  2. Model scores them itself as a judge
  3. Scores become preference pairs for DPO
  4. Trained model becomes next generator and judge
  5. repeat
Self-reward and self-playMeta, NYU (Yuan et al.) · 2024
Claude Code
ProjectAnthropic
Claude Code
  1. Claude Code writes its own code changes
  2. Changes become next Claude Code versions
  3. Release becomes harness for next development round
  4. repeat
Self-modifying agentsAnthropic · 2025
KernelEvolve
ProjectarXiv
KernelEvolve
  1. LLM agents write and benchmark kernels
  2. Kernels speed up AI model training
  3. Session data post-trains smaller kernel models
  4. repeat
Kernels and compilersMeta · 2025
PostTrainBench
BenchmarkarXiv
PostTrainBench
  1. Coding agent writes fine-tuning code
  2. Agent post-trains a base LLM
  3. Trained model scored on target benchmark
  4. One turn of model training another
  5. repeat
AI R&D evalsELLIS Institute Tubingen, MPI-IS, University of Tubingen, Thoughtful Lab (Rank et al.) · 2026
Self-training learned optimizers
FoundationarXiv
Self-training learned optimizers
  1. Learned optimizers update other learned optimizers
  2. Training keeps best performing optimizer parameters
  3. Better optimizers train each other faster
  4. repeat
Architecture and optimizer searchMetz et al. (Google Research, Brain Team) · 2021
CoT monitoring
PaperarXiv
CoT monitoring
  1. LLM reads reasoning model thoughts
  2. LLM flags reward hacking behavior
  3. Signal feeds back into training reward
  4. Trains more aligned capable agents
  5. repeat
Interpretability and oversightOpenAI (Baker et al.) · 2025
Context Language Models
PaperarXiv
Context Language ModelsNew
  1. Model updates context file
  2. Skill-optimization evolves instructions
  3. Model uses improved instructions
  4. Model manages context better
  5. repeat
Inference efficiencyRu-Lin Shao et al. · 2026
AIDE2
PaperarXiv
AIDE2
  1. Outer agent rewrites inner agent code
  2. Accepted rewrites become the incumbent agent
  3. Discovered agent becomes the outer loop agent
  4. repeat
Self-modifying agentsWeco AI · 2026
Redwood
ProjectarXiv
Redwood
  1. AI system writes accelerator hardware
  2. Qwen runs on the new accelerator
  3. Qwen finds optimizations for the accelerator
  4. Optimizations improve next accelerator generation
  5. repeat
Hardware and chipsArchitect Labs · 2026
TPU autoresearch
Mini-projectGitHub
TPU autoresearch
  1. Coding agent edits LLM training code
  2. Agent benchmarks code on TPU hardware
  3. Verdict becomes context for next hypothesis
  4. Wins raise training throughput for same models
  5. repeat
Kernels and compilersAleksey Vlasenko · 2026
Self-optimizing Deep Research
PaperarXiv
Self-optimizing Deep Research
  1. LLM optimizers critique deep research reports
  2. Optimizers rewrite research agent prompts
  3. Optimized system runs the next queries
  4. repeat
Tools and harnessZeta Alpha · 2026
autoresearch-distillation
Mini-projectGitHub
autoresearch-distillation
  1. Qwen3-14B edits GPT training script
  2. Edit is trained and scored
  3. Score rewards Qwen3-14B weights update
  4. Trained checkpoint re-run in autoresearch loop
  5. repeat
Research automationExperiential Labs (Naihin, Fallah) · 2026
FlashInfer-Bench
BenchmarkarXiv
FlashInfer-Bench
  1. LLM agents write GPU kernels
  2. Faster kernels make LLM inference cheaper
  3. New serving traces define next kernel tasks
  4. repeat
Inference efficiencyUW, CMU, NVIDIA, UC Berkeley (Xing et al.) · 2026
Ricursive Intelligence
Projectricursive.com
Ricursive Intelligence
  1. AI systems would design chips
  2. Chips would train more capable AI
  3. That AI designs the next chips
  4. repeat
Hardware and chipsRicursive Intelligence (Anna Goldie, Azalia Mirhoseini) · 2025
Glia
PaperarXiv
Glia
  1. LLM agents write inference cluster code
  2. Designs lower latency and GPU cost
  3. Lowers cost of serving LLMs running Glia
  4. repeat
Inference efficiencyHamadanian et al. (MIT CSAIL) · 2025
ADRS
PaperarXiv
ADRS
  1. LLMs propose and mutate load balancer code
  2. Best programs seed the next generation
  3. Output is faster serving code for LLMs
  4. repeat
Inference efficiencyUC Berkeley (Sky Computing Lab) · 2025
Petri
ProjectAnthropic
Petri
  1. Auditor model tests target model
  2. LLM judge scores target behavior
  3. Feeds into next Claude model release
  4. repeat
Interpretability and oversightAnthropic · 2025
DeepScientist
PaperarXiv
DeepScientist
  1. LLM agents propose hypotheses on AI tasks
  2. Agents implement and test them
  3. Findings Memory steers later proposals
  4. Validated finding directly improves AI system
  5. repeat
Research automationWeng et al. (Westlake University) · 2025
Tongyi DeepResearch
ProjectGitHub
Tongyi DeepResearch
  1. Agent synthesizes agentic pretraining data
  2. Pipeline adjusts training set in real time
  3. Data feeds back into model training
  4. repeat
Training dataTongyi Lab, Alibaba · 2025
ShinkaEvolve
ProjectGitHub
ShinkaEvolve
  1. LLM ensemble rewrites candidate programs
  2. Evaluator scores and selects best programs
  3. Evolved mixture-of-experts load-balancing loss
  4. Loss improves LLM training
  5. repeat
Architecture and optimizer searchSakana AI · 2025
GEPA
ProjectGitHub
GEPA
  1. Reflection LLM writes revised module prompts
  2. Winning prompt candidates are kept
  3. Optimized prompts return to same system
  4. New system traces feed next reflection round
  5. repeat
Tools and harnessAgrawal et al. (UC Berkeley, Stanford, Databricks, MIT) · 2025
LLM Speedrunning Benchmark
BenchmarkarXiv
LLM Speedrunning Benchmark
  1. Agent rewrites LLM training script
  2. Script makes LLM training faster
  3. Benchmark scores the recovered speedup
  4. Speeds up agent's own model training
  5. repeat
AI R&D evalsMeta and collaborators (Zhao et al.) · 2025
Genesys
PaperarXiv
Genesys
  1. Designer agents propose language model architectures
  2. Verifier agents pre-train and evaluate designs
  3. Verified results feed evolutionary population
  4. Language models search for better architectures
  5. repeat
Architecture and optimizer searchCheng, Clark, Richardson (Allen Institute for AI, Dartmouth) · 2025
SEAL
PaperarXiv
SEAL
  1. Language model generates a self-edit
  2. Edit is applied as weight update
  3. RL rewards self-edit by downstream performance
  4. repeat
Training dataMIT (Zweiger, Pari et al.) · 2025
Stanford AI-generated kernels
Mini-projectStanford CRFM
Stanford AI-generated kernels
  1. Frontier LLMs write CUDA kernels
  2. Fastest variants seed each new round
  3. Output feeds next kernel-writing model
  4. repeat
Kernels and compilersStanford CRFM (Ouyang, Mirhoseini, Liang) · 2025
Kevin-32B
ProjectCognition
Kevin-32B
  1. QwQ-32B writes CUDA kernels
  2. Model refines kernels using compiler feedback
  3. RL rewards model on measured speedup
  4. repeat
Kernels and compilersCognition, Stanford · 2025
Absolute Zero
PaperarXiv
Absolute Zero
  1. Model proposes code reasoning tasks
  2. Model solves them for accuracy reward
  3. Python executor verifies the answers
  4. Curriculum shifts as solver improves
  5. repeat
Self-reward and self-playTsinghua, BIGAI, Penn State (Zhao et al.) · 2025
KernelBook + KernelLLM
DatasetHugging Face
KernelBook + KernelLLM
  1. KernelLLM writes Triton GPU kernels
  2. Faster kernels make AI training cheaper
  3. Verified kernels train next kernel model
  4. repeat
Kernels and compilersGPU MODE, Meta · 2025
SWE-smith
Datasetswesmith.com
SWE-smith
  1. o3-mini injects bugs and writes issues
  2. SWE-agent solves tasks for expert trajectories
  3. Trajectories fine-tune Qwen 2.5 Coder
  4. Open coding agent trained on agent data
  5. repeat
Training dataStanford, Princeton, Alibaba Qwen (Yang et al.) · 2025
SICA
PaperarXiv
SICA
  1. Best agent changes its own code
  2. Edited agent is benchmarked and archived
  3. Archived agent makes the next edit
  4. repeat
Self-modifying agentsUniversity of Bristol, iGent AI · 2025
The AI Scientist-v2
ProjectGitHub
The AI Scientist-v2
  1. LLM agents propose ML research ideas
  2. Agents run experiments with tree search
  3. Agents write up results and papers
  4. Output is new knowledge about models
  5. repeat
Research automationYamada et al. (Sakana AI) · 2025
PaperBench
BenchmarkarXiv
PaperBench
  1. Agent replicates published ML research
  2. LLM judge grades the replication
  3. Measures ability to do ML research
  4. Research produces better future models
  5. repeat
AI R&D evalsOpenAI · 2025
rStar-Math
PaperarXiv
rStar-Math
  1. Policy model runs Monte Carlo Tree Search
  2. Verified trajectories train next policy and PPM
  3. PPM guides search to produce better data
  4. Self-evolution improves both models
  5. repeat
Self-reward and self-playMicrosoft Research Asia · 2025
KernelBench
BenchmarkGitHub
KernelBench
  1. LLM writes replacement GPU kernel
  2. Kernel is timed against baseline
  3. Workloads are neural network operators
  4. Makes agent's own training faster
  5. repeat
AI R&D evalsOuyang et al. (Stanford) · 2024
RE-Bench
BenchmarkMETR
RE-Bench
  1. Agents do AI research tasks
  2. Scored against human expert baselines
  3. Measures automation of AI research
  4. Research improves future AI systems
  5. repeat
AI R&D evalsMETR · 2024
MLE-bench
BenchmarkGitHub
MLE-bench
  1. Agents build and train ML models
  2. Submissions scored against human thresholds
  3. Measures open-ended ML research ability
  4. Agents improve their own training code
  5. repeat
AI R&D evalsOpenAI · 2024
ADAS (Meta Agent Search)
PaperarXiv
ADAS (Meta Agent Search)
  1. Meta agent programs new agent designs
  2. Designs are evaluated on target tasks
  3. Results are added to an archive
  4. Meta agent uses archive for next design
  5. repeat
Tools and harnessHu, Lu, Clune (UBC, Vector Institute) · 2024
Self-Taught Evaluator
PaperarXiv
Self-Taught Evaluator
  1. LLM builds worse answers for preference pairs
  2. Judge samples reasoning traces and verdicts
  3. Judge is fine-tuned on correct verdicts
  4. Improved judge labels next iteration's data
  5. repeat
Self-reward and self-playMeta FAIR · 2024
CriticGPT
PaperarXiv
CriticGPT
  1. CriticGPT critiques ChatGPT code answers
  2. Humans catch errors using critiques
  3. Improves labels for RLHF training
  4. Labels train the next ChatGPT
  5. repeat
Interpretability and oversightOpenAI (McAleese et al.) · 2024
DiscoPOP
PaperarXiv
DiscoPOP
  1. GPT-4 proposes preference-optimization losses
  2. Candidate losses fine-tune language models
  3. Evaluation scores feed back into prompt
  4. Best loss aligns other LLMs
  5. repeat
Architecture and optimizer searchSakana AI, FLAIR (Oxford), University of Cambridge · 2024
FineWeb-Edu
DatasetHugging Face
FineWeb-Edu
  1. Llama-3-70B scores pages for educational value
  2. Regressor trained on labels filters corpus
  3. Filtered corpus pretrains new language models
  4. LLM judgment sets next training diet
  5. repeat
Training dataHugging Face · 2024
MAIA
PaperarXiv
MAIA
  1. Agent runs interpretability experiments
  2. Agent finds spurious background cues
  3. Selection retrains classifier final layer
  4. Improves robustness of interpreted model
  5. repeat
Interpretability and oversightMIT CSAIL (Rott Shaham et al.) · 2024
STOP
PaperarXiv
STOP
  1. GPT-4 rewrites the improver program
  2. Each rewrite is scored by meta-utility
  3. Best rewrite becomes next round improver
  4. repeat
Self-modifying agentsZelikman, Lorch, Mackey, Kalai (Stanford, Microsoft Research, OpenAI) · 2023
Aider
ProjectAider
Aider
  1. Aider writes new code for Aider releases
  2. New code becomes the next release
  3. Next release builds the following version
  4. repeat
Self-modifying agentsPaul Gauthier (Aider) · 2023
Model-written evals
PaperarXiv
Model-written evals
  1. Language model writes evaluation questions
  2. Evaluations test other language models
  3. Exposes behaviors worsened by RLHF
  4. Evaluates next round of training
  5. repeat
Interpretability and oversightAnthropic (Perez et al.) · 2022
Constitutional AI
FoundationarXiv
Constitutional AI
  1. Model critiques and revises its own responses
  2. Model is fine-tuned on the revisions
  3. Model judges responses to train preference model
  4. Preference model provides RL reward signal
  5. repeat
Self-reward and self-playBai et al. (Anthropic) · 2022
STaR
FoundationarXiv
STaR
  1. Language model writes step-by-step rationales
  2. Correct rationales are kept as dataset
  3. Model is fine-tuned on self-generated set
  4. Improved model generates next rationale dataset
  5. repeat
Training dataZelikman et al. (Stanford, Google Research) · 2022
AlphaZero
FoundationarXiv
AlphaZero
  1. Network plays games against itself
  2. Games become the training data
  3. Same network trains on its own moves
  4. Updated network plays next self-play games
  5. repeat
Self-reward and self-playSilver et al. (DeepMind) · 2017
NAS with RL
FoundationarXiv
NAS with RL
  1. RNN controller writes child network architecture
  2. Child validation accuracy returns as reward
  3. Updates make controller propose better architectures
  4. repeat
Architecture and optimizer searchZoph and Le (Google Brain) · 2016
Gödel Machine
FoundationarXiv
Gödel Machine
  1. Proof searcher tests proof techniques
  2. Executes useful rewrites of its own code
  3. Rewritten machine searches on its new code
  4. repeat
Self-modifying agentsSchmidhuber (IDSIA) · 2003

Missing a loop?

Send us a project, paper or repo with the loop in one sentence and we will review it for the index.

Submit a project