Alphabell.

AI makes AI better.

A daily hub for work where AI is part of the loop that improves AI, each entry with its loop spelled out. Read more →

1/4
A loop in motion
AlphaEvolve
Project · Google DeepMind
Gemini models propose code changes

Today's radar

5 Oct 2026Full radar →

7 picked from 104 candidates · 31 sources read · 57 loops indexed

  1. 01
    Loop of the dayPaperarXiv2 Oct 2026Training data

    Recursive Self-Improvement in Unified Multimodal Models

    The loop: A unified multimodal model generates images and writes programs to evaluate them, using the verified results to train its own visual understanding and generation.

    Cross-capability self-improvement allows a multimodal model to use one modality to verify and generate training data for another, preventing error accumulation.

    loop fit 10/10Huijuan Wang, Chufan Shi et al.via arXiv
    Recursive Self-Improvement in Unified Multimodal Models
  2. 02
    PaperarXiv2 Oct 2026Tools and harness

    VERSE: Verified Self-Evolving Optimizer for Agent Harnesses

    The loop: The VERSE optimizer tests draft edits and replays failures to revise an agent harness alongside its own prompts and tools, improving its ability to diagnose and fix future errors.

    loop fit 9/10Zekai Wang, Yingqiang Ge et al.via arXiv
    VERSE: Verified Self-Evolving Optimizer for Agent Harnesses
  3. 03
    Forum postLessWrong4 Oct 2026Training data

    Reducing Synthetic Markers Makes Some SDF False Facts Linearly Indistinguishable from Pretraining-Acquired Knowledge

    The loop: A synthetic document generator produces training data with reduced markers, which finetunes a model to internalize false facts that evade middle-layer linear probes.

    loop fit 6/10Jason Zengvia LessWrong
    Reducing Synthetic Markers Makes Some SDF False Facts Linearly Indistinguishable from Pretraining-Acquired Knowledge
  4. 04
    Blog postPyTorch1 Oct 2026Kernels and compilers

    Optimizing Jagged Flash Attention with TLX: The Road Toward SOTA FA4 on Blackwell

    The loop: The Jagged Flash Attention kernel improves inference efficiency on Blackwell, which makes the Generative Ads Model faster.

    loop fit 6/10Han Xu, Jacky Zhou, Jackie (Jiaqi) Xu, Hongtao Yu, Peng Chen (Dev Infra),...via PyTorch
    Optimizing Jagged Flash Attention with TLX: The Road Toward SOTA FA4 on Blackwell
  5. 05
    PaperarXiv2 Oct 2026Training data

    Recursive Harness Self-Improvement for Frontier Reasoning Data Synthesis

    The loop: A data synthesis system recursively improves its own generation harness by converting intermediate solver failures into reusable skills, producing progressively harder reasoning tasks.

    loop fit 9/10Wenlong Zhang, Zhengbo Jiao et al.via arXiv
    Recursive Harness Self-Improvement for Frontier Reasoning Data Synthesis
  6. 06
    Blog postAi21 Oct 2026Tools and harness

    Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

    The loop: Olmo-core 3 provides an open training stack that improves the efficiency of training MoE models, which are then used to build better versions of the stack.

    loop fit 6/10via Ai2
    Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
  7. 07
    PaperarXiv29 Sep 2026Self-reward and self-play

    AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation

    The loop: An advisor model reflects on completed interactions to propose corrections, then uses targeted self-distillation to improve the advice it issues to a frozen language-model executor.

    loop fit 9/10Rishabh Agrawal, Hejie Cui et al.via Hugging Face Papers
    AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation
  8. 08
    PaperarXiv27 Sep 2026Self-modifying agents

    Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents

    The loop: The Rep2Skill agent models its own internal representation trajectories to localize execution errors, generating textual feedback to evolve its external skills.

    loop fit 9/10Euntae Choi, Su-Min Song et al.via Semantic Scholar
    Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents
AlphaEvolve
ProjectGoogle DeepMind
AlphaEvolve
  1. Gemini models propose code changes
  2. Evaluators keep best scoring versions
  3. Finds better Gemini training kernels
  4. Lowers cost to train next Gemini models
  5. repeat
Kernels and compilersGoogle DeepMind · 2025
Darwin Gödel Machine
PaperarXiv
Darwin Gödel Machine
  1. Coding agent edits its own code
  2. Scored child agents enter the archive
  3. Archived agents make later self modifications
  4. repeat
Self-modifying agentsUBC, Vector Institute, Sakana AI · 2025
AlphaChip
ProjectGoogle DeepMind
AlphaChip
  1. RL agent lays out TPU blocks
  2. Better layouts make AI hardware cheaper
  3. Agent pre-trains on earlier chip generations
  4. repeat
Hardware and chipsGoogle DeepMind, Google Research · 2020
autoresearch
Mini-projectGitHub
autoresearch
  1. Coding agent edits LLM training script
  2. Runs a 5-minute training job
  3. Keeps change if validation improves
  4. Changes transferred to larger models
  5. repeat
Research automationAndrej Karpathy · 2026
ScholarEvolve
PaperarXiv
ScholarEvolve
  1. Research agent writes harness mutation blueprints
  2. Coding agent implements harness module changes
  3. Winning harness becomes the new champion
  4. New champion harness runs further generations
  5. repeat
Tools and harnessYang et al. (UC Santa Barbara, Microsoft) · 2026
Nemotron-4 340B
PaperarXiv
Nemotron-4 340B
  1. Aligned model generates training data
  2. Reward model judges and filters it
  3. Data trains next intermediate instruct model
  4. Next model generates better alignment data
  5. repeat
Training dataNVIDIA · 2024
Self-Rewarding LMs
PaperarXiv
Self-Rewarding LMs
  1. Model generates candidate responses to prompts
  2. Model scores them itself as a judge
  3. Scores become preference pairs for DPO
  4. Trained model becomes next generator and judge
  5. repeat
Self-reward and self-playMeta, NYU (Yuan et al.) · 2024
Claude Code
ProjectAnthropic
Claude Code
  1. Claude Code writes its own code changes
  2. Changes become next Claude Code versions
  3. Release becomes harness for next development round
  4. repeat
Self-modifying agentsAnthropic · 2025
KernelEvolve
ProjectarXiv
KernelEvolve
  1. LLM agents write and benchmark kernels
  2. Kernels speed up AI model training
  3. Session data post-trains smaller kernel models
  4. repeat
Kernels and compilersMeta · 2025
PostTrainBench
BenchmarkarXiv
PostTrainBench
  1. Coding agent writes fine-tuning code
  2. Agent post-trains a base LLM
  3. Trained model scored on target benchmark
  4. One turn of model training another
  5. repeat
AI R&D evalsELLIS Institute Tubingen, MPI-IS, University of Tubingen, Thoughtful Lab (Rank et al.) · 2026
Self-training learned optimizers
FoundationarXiv
Self-training learned optimizers
  1. Learned optimizers update other learned optimizers
  2. Training keeps best performing optimizer parameters
  3. Better optimizers train each other faster
  4. repeat
Architecture and optimizer searchMetz et al. (Google Research, Brain Team) · 2021
CoT monitoring
PaperarXiv
CoT monitoring
  1. LLM reads reasoning model thoughts
  2. LLM flags reward hacking behavior
  3. Signal feeds back into training reward
  4. Trains more aligned capable agents
  5. repeat
Interpretability and oversightOpenAI (Baker et al.) · 2025

Support for loop research

How to apply →

Have a loop worth building?

Tell us what the AI improves and how the improvement comes back. Applications for grants, fellowships and compute all start on the contact page.

Get in touch