Alphabell.

AI makes AI better.

A daily hub for work where AI is part of the loop that improves AI, each entry with its loop spelled out. Read more →

1/4
A loop in motion
AlphaEvolve
Project · Google DeepMind
Gemini models propose code changes

Today's radar

11 Oct 2026Full radar →

6 picked from 1 candidates · 31 sources read · 61 loops indexed

  1. 01
    Loop of the dayPaperarXiv7 Oct 2026Self-reward and self-play

    Recurrent Self-Improvement: Dynamic Cross-Loop On-Policy Distillation for Looped Language Models

    The loop: A Looped Language Model uses its own extended recurrent computation as a teacher to supervise its intermediate loops. Because the parameters are shared across all loops, distilling into the intermediate steps updates the shared weights, which immediately improves the terminal loop teacher for the next round of self-improvement.

    By sharing parameters across recurrent steps, LoopOPD elegantly turns a model's own extra compute into a direct, on-policy teacher for its earlier steps, creating a tight and continuous self-improvement loop.

    loop fit 9/10Yi Wang, Rui Qian et al.via arXiv
    Recurrent Self-Improvement: Dynamic Cross-Loop On-Policy Distillation for Looped Language Models
  2. 02
    PaperarXiv8 Oct 2026Self-modifying agents

    Agentic-TTT: Training Test-Time Policy for Test-Time Training

    The loop: A language model learns a test-time policy to decide when and how to apply test-time training algorithms to itself. By training this policy on the utility gains observed from its past decisions, the model becomes better at managing its own parameter-level self-improvement during deployment.

    loop fit 9/10Jiahao Lu, Mohan Kankanhallivia arXiv
    Agentic-TTT: Training Test-Time Policy for Test-Time Training
  3. 03
    PaperarXiv7 Oct 2026Tools and harness

    CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution

    The loop: A terminal agent framework alternates between using execution failures to synthesize a better runtime harness and using verified rollouts to train the model policy. This co-evolution ensures that the model is trained on data matched to its adopted runtime, improving both the agent's weights and the harness it depends on.

    loop fit 9/10Jixuan Chen, Jiaxin Zhang et al.via arXiv
    CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution
  4. 04
    PaperarXiv9 Oct 2026Self-reward and self-play

    EvoRubric: Self-Evolving Rubric-Driven RL for Open-Ended Generation

    The loop: A shared language model policy acts as both a reasoner generating responses and a generator creating evaluation rubrics. Discriminative feedback and peer consensus are used to improve the rubrics, which in turn provide better complementary rewards to optimize the shared policy for open-ended tasks.

    loop fit 9/10Xin Guan, Xiaomeng Hu, Shen Huang, Zhenyi Wang, Bo Zhang, Zijian Li, Pengjun...via arXiv
    EvoRubric: Self-Evolving Rubric-Driven RL for Open-Ended Generation
  5. 05
    PaperarXiv9 Oct 2026Tools and harness

    DuplexAgent-RSI: Recursive Harness Improvement for Full-Duplex Voice Agent Collaboration

    The loop: A full-duplex voice agent system uses interaction traces from simulated conversations to systematically revise its own harness modules. These revisions improve how the system coordinates task delegation and responsiveness, making the agent better at handling complex live collaborations.

    loop fit 9/10Yingda Shen, Yuxiang Wang, Kunyu Feng, Qinke Ni, Jiaqi Li, Minghao Hsu, Junan...via arXiv
    DuplexAgent-RSI: Recursive Harness Improvement for Full-Duplex Voice Agent Collaboration
  6. 06
    PaperarXiv9 Oct 2026Architecture and optimizer search

    Neural Architecture Discovery via Autonomous Evolution

    The loop: The ASI-Arch system autonomously conducts neural architecture research through a closed loop process of experimenting, analyzing, and updating. This discovers novel architectures that can be used to train more capable base models, which in turn power future versions of the research agent.

    loop fit 8/10Weixian Xu, Yixiu Liu, Yang Nan, Lyumanshan Ye, Xiangkun Hu, Zhen Qin, Pengfei...via arXiv
    Neural Architecture Discovery via Autonomous Evolution
  7. 07
    PaperarXiv7 Oct 2026Training data

    AdvSim2Real: Training Web Agents Against Adaptive Prompt Injection in a Web World Model

    The loop: A web world model co-evolves a task curriculum and an injection adversary to train a web agent. The resulting training data improves the agent's robustness and capability, allowing it to handle more difficult tasks and stronger adversaries in the next iteration.

    loop fit 10/10Sarim Hashmi, Mukul Ranjan, Kshitij Mishra, Mikhail Kuznetsov, Praneeth...via arXiv
    AdvSim2Real: Training Web Agents Against Adaptive Prompt Injection in a Web World Model
  8. 08
    PaperarXiv6 Oct 2026Architecture and optimizer search

    RLDiscover: LLM-Driven Co-Evolution of Reinforcement Learning Algorithms

    The loop: The RLDiscover framework progressively co-evolves the components of model-free deep reinforcement learning algorithms. The discovered algorithms improve the learning efficiency of the agents that use them, providing better fitness signals for the framework to discover even stronger algorithms.

    loop fit 9/10Haoran Li, Zengle Ge et al.via arXiv
    RLDiscover: LLM-Driven Co-Evolution of Reinforcement Learning Algorithms
AlphaEvolve
ProjectGoogle DeepMind
AlphaEvolve
  1. Gemini models propose code changes
  2. Evaluators keep best scoring versions
  3. Finds better Gemini training kernels
  4. Lowers cost to train next Gemini models
  5. repeat
Kernels and compilersGoogle DeepMind · 2025
Darwin Gödel Machine
PaperarXiv
Darwin Gödel Machine
  1. Coding agent edits its own code
  2. Scored child agents enter the archive
  3. Archived agents make later self modifications
  4. repeat
Self-modifying agentsUBC, Vector Institute, Sakana AI · 2025
AlphaChip
ProjectGoogle DeepMind
AlphaChip
  1. RL agent lays out TPU blocks
  2. Better layouts make AI hardware cheaper
  3. Agent pre-trains on earlier chip generations
  4. repeat
Hardware and chipsGoogle DeepMind, Google Research · 2020
autoresearch
Mini-projectGitHub
autoresearch
  1. Coding agent edits LLM training script
  2. Runs a 5-minute training job
  3. Keeps change if validation improves
  4. Changes transferred to larger models
  5. repeat
Research automationAndrej Karpathy · 2026
ScholarEvolve
PaperarXiv
ScholarEvolve
  1. Research agent writes harness mutation blueprints
  2. Coding agent implements harness module changes
  3. Winning harness becomes the new champion
  4. New champion harness runs further generations
  5. repeat
Tools and harnessYang et al. (UC Santa Barbara, Microsoft) · 2026
Nemotron-4 340B
PaperarXiv
Nemotron-4 340B
  1. Aligned model generates training data
  2. Reward model judges and filters it
  3. Data trains next intermediate instruct model
  4. Next model generates better alignment data
  5. repeat
Training dataNVIDIA · 2024
Self-Rewarding LMs
PaperarXiv
Self-Rewarding LMs
  1. Model generates candidate responses to prompts
  2. Model scores them itself as a judge
  3. Scores become preference pairs for DPO
  4. Trained model becomes next generator and judge
  5. repeat
Self-reward and self-playMeta, NYU (Yuan et al.) · 2024
Claude Code
ProjectAnthropic
Claude Code
  1. Claude Code writes its own code changes
  2. Changes become next Claude Code versions
  3. Release becomes harness for next development round
  4. repeat
Self-modifying agentsAnthropic · 2025
KernelEvolve
ProjectarXiv
KernelEvolve
  1. LLM agents write and benchmark kernels
  2. Kernels speed up AI model training
  3. Session data post-trains smaller kernel models
  4. repeat
Kernels and compilersMeta · 2025
PostTrainBench
BenchmarkarXiv
PostTrainBench
  1. Coding agent writes fine-tuning code
  2. Agent post-trains a base LLM
  3. Trained model scored on target benchmark
  4. One turn of model training another
  5. repeat
AI R&D evalsELLIS Institute Tubingen, MPI-IS, University of Tubingen, Thoughtful Lab (Rank et al.) · 2026
Self-training learned optimizers
FoundationarXiv
Self-training learned optimizers
  1. Learned optimizers update other learned optimizers
  2. Training keeps best performing optimizer parameters
  3. Better optimizers train each other faster
  4. repeat
Architecture and optimizer searchMetz et al. (Google Research, Brain Team) · 2021
CoT monitoring
PaperarXiv
CoT monitoring
  1. LLM reads reasoning model thoughts
  2. LLM flags reward hacking behavior
  3. Signal feeds back into training reward
  4. Trains more aligned capable agents
  5. repeat
Interpretability and oversightOpenAI (Baker et al.) · 2025

Support for loop research

How to apply →

Have a loop worth building?

Tell us what the AI improves and how the improvement comes back. Applications for grants, fellowships and compute all start on the contact page.

Get in touch