Alphabell.

AI makes AI better.

A daily hub for work where AI is part of the loop that improves AI, each entry with its loop spelled out. Read more →

1/4
A loop in motion
AlphaEvolve
Project · Google DeepMind
Gemini models propose code changes

Today's radar

7 Oct 2026Full radar →

5 picked from 128 candidates · 31 sources read · 58 loops indexed

  1. 01
    Loop of the dayPaperarXiv6 Oct 2026Tools and harness

    RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

    The loop: An LLM agent system iteratively proposes and selects edits to its own harness using a regularized proposer and critic. The resulting improved harness directly upgrades the agent's operating environment, making it better at solving tasks and further improving its harness.

    RRSI demonstrates how regularizing the self-improvement process can prevent an agent from overfitting to its training tasks while evolving its own harness.

    loop fit 9/10Peng Xia, Rujun Han, Zifeng Wang, Yanfei Chen, Yufan Zhuang, Yoonho Lee,...via arXiv
    RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
  2. 02
    PaperarXiv5 Oct 2026Tools and harness

    Second-Order Problem Solving for Recursive Self-Improvement in Formal Verification

    The loop: An agent framework monitors its own execution traces to diagnose structural failures and systematically edit its workflow. These edits improve the agent's workflow for formal verification, making it more effective in subsequent rounds of recursive self-improvement.

    loop fit 9/10Yuxuan Jiang, Aditya Vempaty et al.via arXiv
    Second-Order Problem Solving for Recursive Self-Improvement in Formal Verification
  3. 03
    LeaderboardGitHub Pages6 Oct 2026AI R&D evals

    New best on GSO: claude-fable-5-1_unknown · OpenHands (88.2%)

    The loop: AI agents optimize real codebases for speed. This improves the software infrastructure that AI systems rely on, accelerating the execution of future agents.

    New best on GSO: claude-fable-5-1_unknown · OpenHands (88.2%)
  4. 04
    Forum postLessWrong6 Oct 2026Self-reward and self-play

    Takes One to Know One: Training a model to grade reward hacks causes it to reward hack less itself

    The loop: A grader model is fine-tuned to judge and catch reward hacks in other models. This training simultaneously improves the grader's own alignment, causing it to reward hack less itself.

    loop fit 9/10Arjun Srivia LessWrong
    Takes One to Know One: Training a model to grade reward hacks causes it to reward hack less itself
  5. 05
    PaperarXiv3 Oct 2026Architecture and optimizer search

    EvoCast: Reliable Autonomous Research Agents for Iterative Forecasting Architecture Evolution

    The loop: An autonomous research agent iteratively designs and evaluates time-series forecasting architectures. The experimental outcomes from these evaluations are accumulated as evidence to guide and improve the agent's subsequent architecture search rounds.

    loop fit 9/10Kaipeng Xu, Xianli Yan et al.via arXiv
    EvoCast: Reliable Autonomous Research Agents for Iterative Forecasting Architecture Evolution
  6. 06
    PaperarXiv5 Oct 2026Training data

    Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite

    The loop: A base model discovers successful solutions under diverse harnesses and rewrites them into training trajectories. These trajectories are then used for supervised finetuning, which directly improves the base model's performance on complex tasks.

    loop fit 9/10Zongxia Li, Yucheng Shi, Zhongzhi Li, Junyao Yang, Ruhan Wang, Chengsong...via arXiv
    Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite
  7. 07
    PaperarXiv3 Oct 2026Training data

    Trinity: Self-Evolving Vision-Language Models with a Self-Verifier

    The loop: A vision-language model acts as a Questioner, Solver, and Verifier to generate and screen its own training data. This verified data is then used to train the model's successor, improving its reasoning capabilities without external labels.

    loop fit 9/10Youngwan Lee, Yong-Ju Lee et al.via arXiv
    Trinity: Self-Evolving Vision-Language Models with a Self-Verifier
  8. 08
    PaperarXiv3 Oct 2026Inference efficiency

    SEIS: Self-Evolving Inference Systems

    The loop: SEIS autonomously optimizes its own inference engine code through iterative self-evolution. This redesigns the engine for higher throughput, which accelerates the models that power the system.

    loop fit 10/10Zhen Xu, Jingyu Liu et al.via arXiv
    SEIS: Self-Evolving Inference Systems
AlphaEvolve
ProjectGoogle DeepMind
AlphaEvolve
  1. Gemini models propose code changes
  2. Evaluators keep best scoring versions
  3. Finds better Gemini training kernels
  4. Lowers cost to train next Gemini models
  5. repeat
Kernels and compilersGoogle DeepMind · 2025
Darwin Gödel Machine
PaperarXiv
Darwin Gödel Machine
  1. Coding agent edits its own code
  2. Scored child agents enter the archive
  3. Archived agents make later self modifications
  4. repeat
Self-modifying agentsUBC, Vector Institute, Sakana AI · 2025
AlphaChip
ProjectGoogle DeepMind
AlphaChip
  1. RL agent lays out TPU blocks
  2. Better layouts make AI hardware cheaper
  3. Agent pre-trains on earlier chip generations
  4. repeat
Hardware and chipsGoogle DeepMind, Google Research · 2020
autoresearch
Mini-projectGitHub
autoresearch
  1. Coding agent edits LLM training script
  2. Runs a 5-minute training job
  3. Keeps change if validation improves
  4. Changes transferred to larger models
  5. repeat
Research automationAndrej Karpathy · 2026
ScholarEvolve
PaperarXiv
ScholarEvolve
  1. Research agent writes harness mutation blueprints
  2. Coding agent implements harness module changes
  3. Winning harness becomes the new champion
  4. New champion harness runs further generations
  5. repeat
Tools and harnessYang et al. (UC Santa Barbara, Microsoft) · 2026
Nemotron-4 340B
PaperarXiv
Nemotron-4 340B
  1. Aligned model generates training data
  2. Reward model judges and filters it
  3. Data trains next intermediate instruct model
  4. Next model generates better alignment data
  5. repeat
Training dataNVIDIA · 2024
Self-Rewarding LMs
PaperarXiv
Self-Rewarding LMs
  1. Model generates candidate responses to prompts
  2. Model scores them itself as a judge
  3. Scores become preference pairs for DPO
  4. Trained model becomes next generator and judge
  5. repeat
Self-reward and self-playMeta, NYU (Yuan et al.) · 2024
Claude Code
ProjectAnthropic
Claude Code
  1. Claude Code writes its own code changes
  2. Changes become next Claude Code versions
  3. Release becomes harness for next development round
  4. repeat
Self-modifying agentsAnthropic · 2025
KernelEvolve
ProjectarXiv
KernelEvolve
  1. LLM agents write and benchmark kernels
  2. Kernels speed up AI model training
  3. Session data post-trains smaller kernel models
  4. repeat
Kernels and compilersMeta · 2025
PostTrainBench
BenchmarkarXiv
PostTrainBench
  1. Coding agent writes fine-tuning code
  2. Agent post-trains a base LLM
  3. Trained model scored on target benchmark
  4. One turn of model training another
  5. repeat
AI R&D evalsELLIS Institute Tubingen, MPI-IS, University of Tubingen, Thoughtful Lab (Rank et al.) · 2026
Self-training learned optimizers
FoundationarXiv
Self-training learned optimizers
  1. Learned optimizers update other learned optimizers
  2. Training keeps best performing optimizer parameters
  3. Better optimizers train each other faster
  4. repeat
Architecture and optimizer searchMetz et al. (Google Research, Brain Team) · 2021
CoT monitoring
PaperarXiv
CoT monitoring
  1. LLM reads reasoning model thoughts
  2. LLM flags reward hacking behavior
  3. Signal feeds back into training reward
  4. Trains more aligned capable agents
  5. repeat
Interpretability and oversightOpenAI (Baker et al.) · 2025

Support for loop research

How to apply →

Have a loop worth building?

Tell us what the AI improves and how the improvement comes back. Applications for grants, fellowships and compute all start on the contact page.

Get in touch