Alphabell.

AI makes AI better.

A daily hub for work where AI is part of the loop that improves AI, each entry with its loop spelled out. Read more →

1/4
A loop in motion
AlphaEvolve
Project · Google DeepMind
Gemini models propose code changes

Today's radar

4 Oct 2026Full radar →

7 picked from 5 candidates · 31 sources read · 56 loops indexed

  1. 01
    Loop of the dayPaperarXiv27 Sep 2026Self-modifying agents

    Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents

    The loop: The Rep2Skill agent models its own internal representation trajectories to localize execution errors, generating textual feedback to evolve its external skills.

    Rep2Skill pushes skill evolution beyond text-only feedback by allowing agents to reflect directly on their internal representation trajectories to diagnose and fix execution errors.

    loop fit 9/10Euntae Choi, Su-Min Song et al.via Semantic Scholar
    Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents
  2. 02
    PaperarXiv29 Sep 2026Inference efficiency

    AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

    The loop: A coding agent is trained to decide when and how to compact its own context during long-horizon tasks, using task-success rewards to improve its policy.

    loop fit 9/10Jitin Singla, Parikshit Pareek et al.via arXiv
    AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents
  3. 03
    Blog postNVIDIA28 Sep 2026Hardware and chips

    How NVIDIA DSX MaxLPS Maximizes AI Factory Throughput and Efficiency

    The loop: DSX MaxLPS optimizes power usage in AI factories, which improves throughput for the models running in that factory.

    loop fit 7/10Sarah McKenneyvia NVIDIA Technical Blog
    How NVIDIA DSX MaxLPS Maximizes AI Factory Throughput and Efficiency
  4. 04
    PaperarXiv1 Oct 2026Tools and harness

    LiteEvo: Automated, Cost-Efficient Harness Evolution for Generalization to Unseen Tasks

    The loop: LiteEvo meta-agents mine agent trajectories for reusable components to evolve a harness library, which improves the agent's performance on unseen tasks.

    loop fit 9/10Geyi Yang, Zikun Qu et al.via arXiv
    LiteEvo: Automated, Cost-Efficient Harness Evolution for Generalization to Unseen Tasks
  5. 05
    Blog postPyTorch30 Sep 2026Tools and harness

    From Upstream Changes to Downstream Confidence: Inside Torch Spyre’s Integration with PyTorch CRCR

    The loop: Torch Spyre uses the CRCR relay to automate CI testing, which improves the stability of PyTorch for its own development.

    loop fit 7/10Mehant Kammakomati (IBM), Jewel K M (Red Hat), Anubhav Jana (IBM), Padmanabha...via PyTorch
    From Upstream Changes to Downstream Confidence: Inside Torch Spyre’s Integration with PyTorch CRCR
  6. 06
    PaperarXiv30 Sep 2026Self-reward and self-play

    AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation

    The loop: An advisor model uses reflection to propose corrections to its own decisions, then uses self-distillation from a feedback-conditioned copy to improve its future advice.

    loop fit 9/10Zhijie Wei, Ferris Tan et al.via arXiv
    AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation
  7. 07
    Blog postOpenAI28 Sep 2026Interpretability and oversight

    Towards safety cases for frontier AI training

    The loop: A safety case framework investigates misalignment incidents, which improves technical safeguards for the training of successor frontier models.

    loop fit 6/10via OpenAI
    Towards safety cases for frontier AI training
  8. 08
    PaperarXiv29 Sep 2026Tools and harness

    MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

    The loop: MILO mutator agents rewrite complete agent harnesses and receive parent-specific feedback on their performance. An orchestrator uses this global search history to adapt the mutators' assignments and curriculum, improving the discovery of future harnesses.

    loop fit 9/10Prithwish Jana, Mononito Goswami et al.via arXiv
    MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
AlphaEvolve
ProjectGoogle DeepMind
AlphaEvolve
  1. Gemini models propose code changes
  2. Evaluators keep best scoring versions
  3. Finds better Gemini training kernels
  4. Lowers cost to train next Gemini models
  5. repeat
Kernels and compilersGoogle DeepMind · 2025
Darwin Gödel Machine
PaperarXiv
Darwin Gödel Machine
  1. Coding agent edits its own code
  2. Scored child agents enter the archive
  3. Archived agents make later self modifications
  4. repeat
Self-modifying agentsUBC, Vector Institute, Sakana AI · 2025
AlphaChip
ProjectGoogle DeepMind
AlphaChip
  1. RL agent lays out TPU blocks
  2. Better layouts make AI hardware cheaper
  3. Agent pre-trains on earlier chip generations
  4. repeat
Hardware and chipsGoogle DeepMind, Google Research · 2020
autoresearch
Mini-projectGitHub
autoresearch
  1. Coding agent edits LLM training script
  2. Runs a 5-minute training job
  3. Keeps change if validation improves
  4. Changes transferred to larger models
  5. repeat
Research automationAndrej Karpathy · 2026
ScholarEvolve
PaperarXiv
ScholarEvolve
  1. Research agent writes harness mutation blueprints
  2. Coding agent implements harness module changes
  3. Winning harness becomes the new champion
  4. New champion harness runs further generations
  5. repeat
Tools and harnessYang et al. (UC Santa Barbara, Microsoft) · 2026
Nemotron-4 340B
PaperarXiv
Nemotron-4 340B
  1. Aligned model generates training data
  2. Reward model judges and filters it
  3. Data trains next intermediate instruct model
  4. Next model generates better alignment data
  5. repeat
Training dataNVIDIA · 2024
Self-Rewarding LMs
PaperarXiv
Self-Rewarding LMs
  1. Model generates candidate responses to prompts
  2. Model scores them itself as a judge
  3. Scores become preference pairs for DPO
  4. Trained model becomes next generator and judge
  5. repeat
Self-reward and self-playMeta, NYU (Yuan et al.) · 2024
Claude Code
ProjectAnthropic
Claude Code
  1. Claude Code writes its own code changes
  2. Changes become next Claude Code versions
  3. Release becomes harness for next development round
  4. repeat
Self-modifying agentsAnthropic · 2025
KernelEvolve
ProjectarXiv
KernelEvolve
  1. LLM agents write and benchmark kernels
  2. Kernels speed up AI model training
  3. Session data post-trains smaller kernel models
  4. repeat
Kernels and compilersMeta · 2025
PostTrainBench
BenchmarkarXiv
PostTrainBench
  1. Coding agent writes fine-tuning code
  2. Agent post-trains a base LLM
  3. Trained model scored on target benchmark
  4. One turn of model training another
  5. repeat
AI R&D evalsELLIS Institute Tubingen, MPI-IS, University of Tubingen, Thoughtful Lab (Rank et al.) · 2026
Self-training learned optimizers
FoundationarXiv
Self-training learned optimizers
  1. Learned optimizers update other learned optimizers
  2. Training keeps best performing optimizer parameters
  3. Better optimizers train each other faster
  4. repeat
Architecture and optimizer searchMetz et al. (Google Research, Brain Team) · 2021
CoT monitoring
PaperarXiv
CoT monitoring
  1. LLM reads reasoning model thoughts
  2. LLM flags reward hacking behavior
  3. Signal feeds back into training reward
  4. Trains more aligned capable agents
  5. repeat
Interpretability and oversightOpenAI (Baker et al.) · 2025

Support for loop research

How to apply →

Have a loop worth building?

Tell us what the AI improves and how the improvement comes back. Applications for grants, fellowships and compute all start on the contact page.

Get in touch