Alphabell.

AI makes AI better.

A daily hub for work where AI is part of the loop that improves AI, each entry with its loop spelled out. Read more →

1/4
A loop in motion
AlphaEvolve
Project · Google DeepMind
Gemini models propose code changes

Today's radar

3 Oct 2026Full radar →

8 picked from 3 candidates · 31 sources read · 56 loops indexed

  1. 01
    Loop of the dayPaperarXiv29 Sep 2026Tools and harness

    MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

    The loop: MILO mutator agents rewrite complete agent harnesses and receive parent-specific feedback on their performance. An orchestrator uses this global search history to adapt the mutators' assignments and curriculum, improving the discovery of future harnesses.

    MILO shows how co-evolving an agent's harness alongside the strategy used to discover it can break out of fixed search patterns and yield superior execution environments.

    loop fit 9/10Prithwish Jana, Mononito Goswami et al.via arXiv
    MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
  2. 02
    PaperarXiv29 Sep 2026Self-modifying agents

    Topological Coherence for Self-evolving Multi-agent Systems

    The loop: TOCOMAS proposes coupled changes to its own agent, collaboration, and memory policies during online execution. It retains candidates that satisfy structural constraints and improve evaluated reward, directly evolving its own architecture for future tasks.

    loop fit 9/10Sen Zhao, Ruiqi Kong et al.via arXiv
    Topological Coherence for Self-evolving Multi-agent Systems
  3. 03
    ProjectGitHub3 Mar 2026Tools and harness

    IgorGanapolsky/ThumbGate

    The loop: ThumbGate Pre-Action Checks analyze ranked lessons and repeated failures from past executions. The system uses this data to self-improve its strict mode blocking rules, becoming better at preventing secret leaks in future actions.

    loop fit 7/10IgorGanapolskyvia GitHub
    IgorGanapolsky/ThumbGate
  4. 04
    ModelHugging Face2 Oct 2026Training data

    AutoSynthData: Generating Training Data for Enterprise Agents

    The loop: The AutoSynthData generator creates synthetic training datasets tailored for enterprise environments. This data is then fed back into the training pipeline to improve the performance and reliability of the enterprise agents.

    loop fit 6/10via Hugging Face
    AutoSynthData: Generating Training Data for Enterprise Agents
  5. 05
    PaperarXiv29 Sep 2026Self-reward and self-play

    Train Ahead, Distill Back: Bootstrapping On-Policy Self-Distillation for Large Language Models

    The loop: The B-OPSD policy temporarily trains ahead to create a stronger future teacher. This teacher then generates reliable trajectories and provides dense supervision to the restarted original student, improving the model's own successor.

    loop fit 9/10Zheng Zhang, Xinyue Tan et al.via arXiv
    Train Ahead, Distill Back: Bootstrapping On-Policy Self-Distillation for Large Language Models
  6. 06
    ProjectGitHub17 Aug 2026Tools and harness

    proteus-evolve/Proteus

    The loop: Proteus plugs into any agent harness to measure its performance and propose evolutionary changes. These changes are applied to the harness, improving the execution environment for subsequent agent runs.

    loop fit 6/10proteus-evolvevia GitHub
    proteus-evolve/Proteus
  7. 07
    PaperarXiv26 Sep 2026Training data

    X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization

    The loop: The X-Tree tokenizer mines flat action streams to build a hierarchy of reusable skills without LLM calls. This tree is then used as a self-teacher during on-policy self-distillation, improving the agent's ability to generalize across tasks.

    loop fit 9/10Sitao Cheng, Xunjian Yin et al.via Hugging Face Papers
    X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization
  8. 08
    ProjectGitHub13 Aug 2026Tools and harness

    ruvnet/dream-machine

    The loop: The Dream Machine engine schedules nightly repository evolution tasks and evaluates the results. It uses an evidence-gated promotion system to merge successful changes, continuously improving its own configuration and codebase.

    loop fit 6/10ruvnetvia GitHub
    ruvnet/dream-machine
AlphaEvolve
ProjectGoogle DeepMind
AlphaEvolve
  1. Gemini models propose code changes
  2. Evaluators keep best scoring versions
  3. Finds better Gemini training kernels
  4. Lowers cost to train next Gemini models
  5. repeat
Kernels and compilersGoogle DeepMind · 2025
Darwin Gödel Machine
PaperarXiv
Darwin Gödel Machine
  1. Coding agent edits its own code
  2. Scored child agents enter the archive
  3. Archived agents make later self modifications
  4. repeat
Self-modifying agentsUBC, Vector Institute, Sakana AI · 2025
AlphaChip
ProjectGoogle DeepMind
AlphaChip
  1. RL agent lays out TPU blocks
  2. Better layouts make AI hardware cheaper
  3. Agent pre-trains on earlier chip generations
  4. repeat
Hardware and chipsGoogle DeepMind, Google Research · 2020
autoresearch
Mini-projectGitHub
autoresearch
  1. Coding agent edits LLM training script
  2. Runs a 5-minute training job
  3. Keeps change if validation improves
  4. Changes transferred to larger models
  5. repeat
Research automationAndrej Karpathy · 2026
ScholarEvolve
PaperarXiv
ScholarEvolve
  1. Research agent writes harness mutation blueprints
  2. Coding agent implements harness module changes
  3. Winning harness becomes the new champion
  4. New champion harness runs further generations
  5. repeat
Tools and harnessYang et al. (UC Santa Barbara, Microsoft) · 2026
Nemotron-4 340B
PaperarXiv
Nemotron-4 340B
  1. Aligned model generates training data
  2. Reward model judges and filters it
  3. Data trains next intermediate instruct model
  4. Next model generates better alignment data
  5. repeat
Training dataNVIDIA · 2024
Self-Rewarding LMs
PaperarXiv
Self-Rewarding LMs
  1. Model generates candidate responses to prompts
  2. Model scores them itself as a judge
  3. Scores become preference pairs for DPO
  4. Trained model becomes next generator and judge
  5. repeat
Self-reward and self-playMeta, NYU (Yuan et al.) · 2024
Claude Code
ProjectAnthropic
Claude Code
  1. Claude Code writes its own code changes
  2. Changes become next Claude Code versions
  3. Release becomes harness for next development round
  4. repeat
Self-modifying agentsAnthropic · 2025
KernelEvolve
ProjectarXiv
KernelEvolve
  1. LLM agents write and benchmark kernels
  2. Kernels speed up AI model training
  3. Session data post-trains smaller kernel models
  4. repeat
Kernels and compilersMeta · 2025
PostTrainBench
BenchmarkarXiv
PostTrainBench
  1. Coding agent writes fine-tuning code
  2. Agent post-trains a base LLM
  3. Trained model scored on target benchmark
  4. One turn of model training another
  5. repeat
AI R&D evalsELLIS Institute Tubingen, MPI-IS, University of Tubingen, Thoughtful Lab (Rank et al.) · 2026
Self-training learned optimizers
FoundationarXiv
Self-training learned optimizers
  1. Learned optimizers update other learned optimizers
  2. Training keeps best performing optimizer parameters
  3. Better optimizers train each other faster
  4. repeat
Architecture and optimizer searchMetz et al. (Google Research, Brain Team) · 2021
CoT monitoring
PaperarXiv
CoT monitoring
  1. LLM reads reasoning model thoughts
  2. LLM flags reward hacking behavior
  3. Signal feeds back into training reward
  4. Trains more aligned capable agents
  5. repeat
Interpretability and oversightOpenAI (Baker et al.) · 2025

Support for loop research

How to apply →

Have a loop worth building?

Tell us what the AI improves and how the improvement comes back. Applications for grants, fellowships and compute all start on the contact page.

Get in touch