Alphabell.

AI makes AI better.

A daily hub for work where AI is part of the loop that improves AI, each entry with its loop spelled out. Read more →

1/4
A loop in motion
AlphaEvolve
Project · Google DeepMind
Gemini models propose code changes

Today's radar

9 Oct 2026Full radar →

9 picked from 134 candidates · 31 sources read · 60 loops indexed

  1. 01
    Loop of the dayPaperarXiv9 Oct 2026Self-reward and self-play

    The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators

    The loop: A coder agent and a reviewer agent co-evolve under non-stationary utilities. The reviewer grades patches to guide the coder's search, while the coder's output helps the reviewer refine its grading rubric, making both better at their respective tasks.

    By co-evolving the coder and its evaluator, the Red Queen Gödel Machine escapes the bottleneck of static reward functions, allowing the evaluation criteria to scale alongside the agent's capabilities.

    loop fit 10/10Alex Iacob, Andrej Jovanovi\'c, William F. Shen, Daniel Burkhardt, Meghdad...via arXiv
    The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
  2. 02
    PaperarXiv7 Oct 2026Kernels and compilers

    KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization

    The loop: A multi-agent system profiles and optimizes generated Triton sub-kernels within compiled models. This improves the execution efficiency of the models and allows the agents to verify the re-stitched model end-to-end.

    loop fit 9/10Aheli Poddar, Sanskar Prasad, Arindam Samanta, Subha Chakraborty, Vishal...via arXiv
    KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization
  3. 03
    ProjectGitHub24 Jun 2026Tools and harness

    Ker102/Harneloop

    The loop: Self-improving AI agents use trace-backed diagnosis and artifact-aware testing to refine their own harness units. This creates a better execution environment for their subsequent runs.

    loop fit 7/10Ker102via GitHub
    Ker102/Harneloop
  4. 04
    Forum postLessWrong8 Oct 2026Interpretability and oversight

    J++ Lens: Jacobian Filtering Enables More Faithful Workspace Lenses

    The loop: The J++ Lens provides more faithful readouts of model activations, which improves the ability to monitor and oversee the training of future models.

    loop fit 7/10Kola Ayonrindevia LessWrong
    J++ Lens: Jacobian Filtering Enables More Faithful Workspace Lenses
  5. 05
    Blog postAnthropic8 Oct 2026Tools and harness

    An opt-in vulnerability-finding service for open-source software

    The loop: Claude scans open source software for vulnerabilities, which improves the security of the libraries Claude uses, making the system more robust.

    loop fit 6/10via Anthropic
    An opt-in vulnerability-finding service for open-source software
  6. 06
    PaperarXiv8 Oct 2026Research automation

    Language Models as AI Research World Models

    The loop: Language models act as Research World Models to predict the outcomes of candidate AI research experiments. This improves the selection of future interventions and generates new experimental data to further train the world model.

    loop fit 9/10Zijun Wang, Zewen Liu et al.via arXiv
    Language Models as AI Research World Models
  7. 07
    ProjectGitHub6 Oct 2026Research automation

    richardcsuwandi/kernaut

    The loop: An automated research agent searches for novel GPU kernels. These kernels can then be used to discover and train more efficient architectures for future versions of the agent.

    loop fit 6/10richardcsuwandivia GitHub
    richardcsuwandi/kernaut
  8. 08
    PaperarXiv7 Oct 2026AI R&D evals

    RSIGym: A Flexible Environment for Recursive Self-Improvement

    The loop: AI research agents use a flexible environment to propose and evaluate changes to their own training data and execution harnesses. Successful interventions are then carried forward into subsequent improvement cycles, making the agents better at future research tasks.

    loop fit 9/10Fanqing Meng, Lingxiao Du et al.via arXiv
    RSIGym: A Flexible Environment for Recursive Self-Improvement
AlphaEvolve
ProjectGoogle DeepMind
AlphaEvolve
  1. Gemini models propose code changes
  2. Evaluators keep best scoring versions
  3. Finds better Gemini training kernels
  4. Lowers cost to train next Gemini models
  5. repeat
Kernels and compilersGoogle DeepMind · 2025
Darwin Gödel Machine
PaperarXiv
Darwin Gödel Machine
  1. Coding agent edits its own code
  2. Scored child agents enter the archive
  3. Archived agents make later self modifications
  4. repeat
Self-modifying agentsUBC, Vector Institute, Sakana AI · 2025
AlphaChip
ProjectGoogle DeepMind
AlphaChip
  1. RL agent lays out TPU blocks
  2. Better layouts make AI hardware cheaper
  3. Agent pre-trains on earlier chip generations
  4. repeat
Hardware and chipsGoogle DeepMind, Google Research · 2020
autoresearch
Mini-projectGitHub
autoresearch
  1. Coding agent edits LLM training script
  2. Runs a 5-minute training job
  3. Keeps change if validation improves
  4. Changes transferred to larger models
  5. repeat
Research automationAndrej Karpathy · 2026
ScholarEvolve
PaperarXiv
ScholarEvolve
  1. Research agent writes harness mutation blueprints
  2. Coding agent implements harness module changes
  3. Winning harness becomes the new champion
  4. New champion harness runs further generations
  5. repeat
Tools and harnessYang et al. (UC Santa Barbara, Microsoft) · 2026
Nemotron-4 340B
PaperarXiv
Nemotron-4 340B
  1. Aligned model generates training data
  2. Reward model judges and filters it
  3. Data trains next intermediate instruct model
  4. Next model generates better alignment data
  5. repeat
Training dataNVIDIA · 2024
Self-Rewarding LMs
PaperarXiv
Self-Rewarding LMs
  1. Model generates candidate responses to prompts
  2. Model scores them itself as a judge
  3. Scores become preference pairs for DPO
  4. Trained model becomes next generator and judge
  5. repeat
Self-reward and self-playMeta, NYU (Yuan et al.) · 2024
Claude Code
ProjectAnthropic
Claude Code
  1. Claude Code writes its own code changes
  2. Changes become next Claude Code versions
  3. Release becomes harness for next development round
  4. repeat
Self-modifying agentsAnthropic · 2025
KernelEvolve
ProjectarXiv
KernelEvolve
  1. LLM agents write and benchmark kernels
  2. Kernels speed up AI model training
  3. Session data post-trains smaller kernel models
  4. repeat
Kernels and compilersMeta · 2025
PostTrainBench
BenchmarkarXiv
PostTrainBench
  1. Coding agent writes fine-tuning code
  2. Agent post-trains a base LLM
  3. Trained model scored on target benchmark
  4. One turn of model training another
  5. repeat
AI R&D evalsELLIS Institute Tubingen, MPI-IS, University of Tubingen, Thoughtful Lab (Rank et al.) · 2026
Self-training learned optimizers
FoundationarXiv
Self-training learned optimizers
  1. Learned optimizers update other learned optimizers
  2. Training keeps best performing optimizer parameters
  3. Better optimizers train each other faster
  4. repeat
Architecture and optimizer searchMetz et al. (Google Research, Brain Team) · 2021
CoT monitoring
PaperarXiv
CoT monitoring
  1. LLM reads reasoning model thoughts
  2. LLM flags reward hacking behavior
  3. Signal feeds back into training reward
  4. Trains more aligned capable agents
  5. repeat
Interpretability and oversightOpenAI (Baker et al.) · 2025

Support for loop research

How to apply →

Have a loop worth building?

Tell us what the AI improves and how the improvement comes back. Applications for grants, fellowships and compute all start on the contact page.

Get in touch