Alphabell.

AI makes AI better.

A daily hub for work where AI is part of the loop that improves AI, each entry with its loop spelled out. Read more →

1/4
A loop in motion
AlphaEvolve
Project · Google DeepMind
Gemini models propose code changes

Today's radar

2 Oct 2026Full radar →

8 picked from 56 candidates · 31 sources read · 56 loops indexed

  1. 01
    Loop of the dayPaperarXiv29 Sep 2026Training data

    AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks

    The loop: The AREX-2 agent synthesizes long-horizon improvement trajectories from ML engineering tasks and uses them to train its successor for better performance on MLE-bench.

    AREX-2 demonstrates that long-horizon reflective data synthesized from ML engineering tasks can effectively train agents to become better at self-improvement.

    loop fit 9/10Hongjin Qian, Chaofan Li et al.via arXiv
    AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks
  2. 02
    PaperarXiv29 Sep 2026Self-modifying agents

    SelfSearch: Reward-Free Search for Self-Improving Agents

    The loop: An LLM agent modifies its own instructions and tools using records of its own previous attempts, improving its success rate on future tasks.

    loop fit 9/10Jungwoo Yang, Injin Kong et al.via arXiv
    SelfSearch: Reward-Free Search for Self-Improving Agents
  3. 03
    ProjectGitHub17 Apr 2026Self-modifying agents

    ModernOps888/the-forge

    The loop: Four LLMs compete to evolve and breed code through a JIT compiler judge, creating a closed loop of code improvement that enhances their own capabilities.

    loop fit 9/10ModernOps888via GitHub
    ModernOps888/the-forge
  4. 04
    ReleaseGitHub28 Sep 2026Self-modifying agents

    algorithmicsuperintelligence/openevolve: v0.4.0

    The loop: The OpenEvolve system uses pluggable selection strategies and program validation to evolve code, capturing token usage and enforcing evolution blocks for its own mutations.

    loop fit 8/10codelionvia GitHub releases
    algorithmicsuperintelligence/openevolve: v0.4.0
  5. 05
    Blog postPyTorch2 Oct 2026Kernels and compilers

    Building a High-Performance and Portable vLLM Linear Backend with Helion

    The loop: The Helion autotuner uses a DSL to generate high-performance kernels for vLLM, which then runs LLM inference more efficiently for future tasks.

    loop fit 7/10Sean Chen (Red Hat) and Shangdi Yu (PyTorch, Meta Platforms)via PyTorch
    Building a High-Performance and Portable vLLM Linear Backend with Helion
  6. 06
    PaperarXiv29 Sep 2026Self-reward and self-play

    RLTL;DR: Self-improvement by Internalizing Self-generated Feedback

    The loop: A policy model generates its own feedback insights from failed attempts and internalizes them through training, improving its success rate on future rollouts.

    loop fit 9/10Michael Kirchhof, Eleonora Gualdoni et al.via arXiv
    RLTL;DR: Self-improvement by Internalizing Self-generated Feedback
  7. 07
    ProjectGitHub13 Aug 2026Tools and harness

    ZK-Andy/dsh-continual-evolve

    The loop: A continual self-evolution plugin refines the DeepSeek Harness state based on session trajectories, using a benchmark-driven loop to validate and improve its own harness.

    loop fit 8/10ZK-Andyvia GitHub
    ZK-Andy/dsh-continual-evolve
  8. 08
    PaperarXiv29 Sep 2026Tools and harness

    Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer

    The loop: A frozen model acts as a solver to generate run records and then as a proposer to edit its own harness, improving its performance on subsequent tasks.

    loop fit 9/10Qiankai Xuvia arXiv
    Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer
AlphaEvolve
ProjectGoogle DeepMind
AlphaEvolve
  1. Gemini models propose code changes
  2. Evaluators keep best scoring versions
  3. Finds better Gemini training kernels
  4. Lowers cost to train next Gemini models
  5. repeat
Kernels and compilersGoogle DeepMind · 2025
Darwin Gödel Machine
PaperarXiv
Darwin Gödel Machine
  1. Coding agent edits its own code
  2. Scored child agents enter the archive
  3. Archived agents make later self modifications
  4. repeat
Self-modifying agentsUBC, Vector Institute, Sakana AI · 2025
AlphaChip
ProjectGoogle DeepMind
AlphaChip
  1. RL agent lays out TPU blocks
  2. Better layouts make AI hardware cheaper
  3. Agent pre-trains on earlier chip generations
  4. repeat
Hardware and chipsGoogle DeepMind, Google Research · 2020
autoresearch
Mini-projectGitHub
autoresearch
  1. Coding agent edits LLM training script
  2. Runs a 5-minute training job
  3. Keeps change if validation improves
  4. Changes transferred to larger models
  5. repeat
Research automationAndrej Karpathy · 2026
ScholarEvolve
PaperarXiv
ScholarEvolve
  1. Research agent writes harness mutation blueprints
  2. Coding agent implements harness module changes
  3. Winning harness becomes the new champion
  4. New champion harness runs further generations
  5. repeat
Tools and harnessYang et al. (UC Santa Barbara, Microsoft) · 2026
Nemotron-4 340B
PaperarXiv
Nemotron-4 340B
  1. Aligned model generates training data
  2. Reward model judges and filters it
  3. Data trains next intermediate instruct model
  4. Next model generates better alignment data
  5. repeat
Training dataNVIDIA · 2024
Self-Rewarding LMs
PaperarXiv
Self-Rewarding LMs
  1. Model generates candidate responses to prompts
  2. Model scores them itself as a judge
  3. Scores become preference pairs for DPO
  4. Trained model becomes next generator and judge
  5. repeat
Self-reward and self-playMeta, NYU (Yuan et al.) · 2024
Claude Code
ProjectAnthropic
Claude Code
  1. Claude Code writes its own code changes
  2. Changes become next Claude Code versions
  3. Release becomes harness for next development round
  4. repeat
Self-modifying agentsAnthropic · 2025
KernelEvolve
ProjectarXiv
KernelEvolve
  1. LLM agents write and benchmark kernels
  2. Kernels speed up AI model training
  3. Session data post-trains smaller kernel models
  4. repeat
Kernels and compilersMeta · 2025
PostTrainBench
BenchmarkarXiv
PostTrainBench
  1. Coding agent writes fine-tuning code
  2. Agent post-trains a base LLM
  3. Trained model scored on target benchmark
  4. One turn of model training another
  5. repeat
AI R&D evalsELLIS Institute Tubingen, MPI-IS, University of Tubingen, Thoughtful Lab (Rank et al.) · 2026
Self-training learned optimizers
FoundationarXiv
Self-training learned optimizers
  1. Learned optimizers update other learned optimizers
  2. Training keeps best performing optimizer parameters
  3. Better optimizers train each other faster
  4. repeat
Architecture and optimizer searchMetz et al. (Google Research, Brain Team) · 2021
CoT monitoring
PaperarXiv
CoT monitoring
  1. LLM reads reasoning model thoughts
  2. LLM flags reward hacking behavior
  3. Signal feeds back into training reward
  4. Trains more aligned capable agents
  5. repeat
Interpretability and oversightOpenAI (Baker et al.) · 2025

Support for loop research

How to apply →

Have a loop worth building?

Tell us what the AI improves and how the improvement comes back. Applications for grants, fellowships and compute all start on the contact page.

Get in touch