Alphabell.

AI makes AI better.

A daily hub for work where AI is part of the loop that improves AI, each entry with its loop spelled out. Read more →

1/4
A loop in motion
AlphaEvolve
Project · Google DeepMind
Gemini models propose code changes

Today's radar

6 Oct 2026Full radar →

6 picked from 122 candidates · 31 sources read · 58 loops indexed

  1. 01
    Loop of the dayPaperarXiv6 Oct 2026Hardware and chips

    SparseCraft: Agentic Hardware-Software Co-Optimization for Sparse Computing

    The loop: A language model iteratively edits the RTL and schedules of a sparse accelerator based on measured simulation outcomes. This produces a faster and more efficient hardware design, which can then run the model itself more efficiently.

    SparseCraft demonstrates a closed loop where a language model optimizes the RTL and schedules of a hardware accelerator, paving the way for models to design the chips that run them.

    loop fit 10/10Rajatabha Chakraborty, M P Samartha, Vedant Pahariya, Priyesh Shuklavia arXiv
    SparseCraft: Agentic Hardware-Software Co-Optimization for Sparse Computing
  2. 02
    PaperarXiv3 Oct 2026Inference efficiency

    SEIS: Self-Evolving Inference Systems

    The loop: SEIS autonomously optimizes its own inference engine code through iterative self-evolution. This redesigns the engine for higher throughput, which accelerates the models that power the system.

    loop fit 10/10Zhen Xu, Jingyu Liu et al.via arXiv
    SEIS: Self-Evolving Inference Systems
  3. 03
    LeaderboardGitHub Pages6 Oct 2026AI R&D evals

    New best on GSO: claude-fable-5-1_unknown · OpenHands (88.2%)

    The loop: AI agents optimize real codebases for speed. This improves the software infrastructure that AI systems rely on, accelerating the execution of future agents.

    New best on GSO: claude-fable-5-1_unknown · OpenHands (88.2%)
  4. 04
    Forum postLessWrong6 Oct 2026Self-reward and self-play

    Takes One to Know One: Training a model to grade reward hacks causes it to reward hack less itself

    The loop: A grader model is fine-tuned to judge and catch reward hacks in other models. This training simultaneously improves the grader's own alignment, causing it to reward hack less itself.

    loop fit 9/10Arjun Srivia LessWrong
    Takes One to Know One: Training a model to grade reward hacks causes it to reward hack less itself
  5. 05
    Blog postEpoch AI2 Oct 2026Research automation

    Coding-agent use at OpenAI is doubling roughly every month

    The loop: Coding agents assist OpenAI researchers in automating the development of AI software. This accelerates the engineering of the next generation of models and agents.

    loop fit 6/10via Epoch AI
    Coding-agent use at OpenAI is doubling roughly every month
  6. 06
    PaperarXiv4 Oct 2026Tools and harness

    MESH-Harness: Self-Improving Agent Harnesses via Bandit-Guided Compositional Evolution

    The loop: An agent uses bandit-guided composition and local code edits to iteratively evolve its own harness modules. This produces a better harness that improves the agent's performance on future tasks.

    loop fit 9/10Zhiwei Shang, Yu Huo et al.via arXiv
    MESH-Harness: Self-Improving Agent Harnesses via Bandit-Guided Compositional Evolution
  7. 07
    PaperarXiv2 Oct 2026Training data

    Recursive Self-Improvement in Unified Multimodal Models

    The loop: A unified multimodal model generates images and writes programs to evaluate them, using the verified results to train its own visual understanding and generation.

    loop fit 10/10Huijuan Wang, Chufan Shi et al.via arXiv
    Recursive Self-Improvement in Unified Multimodal Models
  8. 08
    PaperarXiv2 Oct 2026Tools and harness

    VERSE: Verified Self-Evolving Optimizer for Agent Harnesses

    The loop: The VERSE optimizer tests draft edits and replays failures to revise an agent harness alongside its own prompts and tools, improving its ability to diagnose and fix future errors.

    loop fit 9/10Zekai Wang, Yingqiang Ge et al.via arXiv
    VERSE: Verified Self-Evolving Optimizer for Agent Harnesses
AlphaEvolve
ProjectGoogle DeepMind
AlphaEvolve
  1. Gemini models propose code changes
  2. Evaluators keep best scoring versions
  3. Finds better Gemini training kernels
  4. Lowers cost to train next Gemini models
  5. repeat
Kernels and compilersGoogle DeepMind · 2025
Darwin Gödel Machine
PaperarXiv
Darwin Gödel Machine
  1. Coding agent edits its own code
  2. Scored child agents enter the archive
  3. Archived agents make later self modifications
  4. repeat
Self-modifying agentsUBC, Vector Institute, Sakana AI · 2025
AlphaChip
ProjectGoogle DeepMind
AlphaChip
  1. RL agent lays out TPU blocks
  2. Better layouts make AI hardware cheaper
  3. Agent pre-trains on earlier chip generations
  4. repeat
Hardware and chipsGoogle DeepMind, Google Research · 2020
autoresearch
Mini-projectGitHub
autoresearch
  1. Coding agent edits LLM training script
  2. Runs a 5-minute training job
  3. Keeps change if validation improves
  4. Changes transferred to larger models
  5. repeat
Research automationAndrej Karpathy · 2026
ScholarEvolve
PaperarXiv
ScholarEvolve
  1. Research agent writes harness mutation blueprints
  2. Coding agent implements harness module changes
  3. Winning harness becomes the new champion
  4. New champion harness runs further generations
  5. repeat
Tools and harnessYang et al. (UC Santa Barbara, Microsoft) · 2026
Nemotron-4 340B
PaperarXiv
Nemotron-4 340B
  1. Aligned model generates training data
  2. Reward model judges and filters it
  3. Data trains next intermediate instruct model
  4. Next model generates better alignment data
  5. repeat
Training dataNVIDIA · 2024
Self-Rewarding LMs
PaperarXiv
Self-Rewarding LMs
  1. Model generates candidate responses to prompts
  2. Model scores them itself as a judge
  3. Scores become preference pairs for DPO
  4. Trained model becomes next generator and judge
  5. repeat
Self-reward and self-playMeta, NYU (Yuan et al.) · 2024
Claude Code
ProjectAnthropic
Claude Code
  1. Claude Code writes its own code changes
  2. Changes become next Claude Code versions
  3. Release becomes harness for next development round
  4. repeat
Self-modifying agentsAnthropic · 2025
KernelEvolve
ProjectarXiv
KernelEvolve
  1. LLM agents write and benchmark kernels
  2. Kernels speed up AI model training
  3. Session data post-trains smaller kernel models
  4. repeat
Kernels and compilersMeta · 2025
PostTrainBench
BenchmarkarXiv
PostTrainBench
  1. Coding agent writes fine-tuning code
  2. Agent post-trains a base LLM
  3. Trained model scored on target benchmark
  4. One turn of model training another
  5. repeat
AI R&D evalsELLIS Institute Tubingen, MPI-IS, University of Tubingen, Thoughtful Lab (Rank et al.) · 2026
Self-training learned optimizers
FoundationarXiv
Self-training learned optimizers
  1. Learned optimizers update other learned optimizers
  2. Training keeps best performing optimizer parameters
  3. Better optimizers train each other faster
  4. repeat
Architecture and optimizer searchMetz et al. (Google Research, Brain Team) · 2021
CoT monitoring
PaperarXiv
CoT monitoring
  1. LLM reads reasoning model thoughts
  2. LLM flags reward hacking behavior
  3. Signal feeds back into training reward
  4. Trains more aligned capable agents
  5. repeat
Interpretability and oversightOpenAI (Baker et al.) · 2025

Support for loop research

How to apply →

Have a loop worth building?

Tell us what the AI improves and how the improvement comes back. Applications for grants, fellowships and compute all start on the contact page.

Get in touch