Alphabell.
Daily edition

Radar, 9 Oct 2026

Today's edition highlights the shift from static evaluation to co-evolution, featuring a Gödel machine that evolves its own evaluators alongside its agents.

PaperarXiv·9 Oct 2026·Self-reward and self-play

The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators

The loop: A coder agent and a reviewer agent co-evolve under non-stationary utilities. The reviewer grades patches to guide the coder's search, while the coder's output helps the reviewer refine its grading rubric, making both better at their respective tasks.

An evolutionary framework that enables recursive self-improvement by co-evolving agents alongside the evaluators that guide their search.

loop fit 10/10Alex Iacob, Andrej Jovanovi\'c, William F. Shen, Daniel Burkhardt, Meghdad...via arXiv
The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
PaperarXiv·7 Oct 2026·Kernels and compilers

KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization

The loop: A multi-agent system profiles and optimizes generated Triton sub-kernels within compiled models. This improves the execution efficiency of the models and allows the agents to verify the re-stitched model end-to-end.

A multi-agent system that optimizes GPU kernels by treating compiled models as structured artifacts and exclusively targeting generated Triton sub-kernels.

loop fit 9/10Aheli Poddar, Sanskar Prasad, Arindam Samanta, Subha Chakraborty, Vishal...via arXiv
KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization
PaperarXiv·8 Oct 2026·Research automation

Language Models as AI Research World Models

The loop: Language models act as Research World Models to predict the outcomes of candidate AI research experiments. This improves the selection of future interventions and generates new experimental data to further train the world model.

An investigation into using language models to predict the outcomes of candidate interventions across AI research environments to sustain self-improvement under limited budgets.

loop fit 9/10Zijun Wang, Zewen Liu, Minhua Lin et al.via arXiv
Language Models as AI Research World Models
PaperarXiv·7 Oct 2026·AI R&D evals

RSIGym: A Flexible Environment for Recursive Self-Improvement

The loop: AI research agents use a flexible environment to propose and evaluate changes to their own training data and execution harnesses. Successful interventions are then carried forward into subsequent improvement cycles, making the agents better at future research tasks.

An agent-native research environment that exposes training, inference, and evaluation through reusable services to support recursive self-improvement.

loop fit 9/10Fanqing Meng, Lingxiao Du, Haocheng Lu et al.via arXiv
RSIGym: A Flexible Environment for Recursive Self-Improvement
PaperarXiv·6 Oct 2026·Self-modifying agents

Co-Evolving Robot Orchestrators and Policies through Deployment

The loop: A vision-language model orchestrator curates skill demonstrations from its own executions to fine-tune its underlying policy. This removes the policy bottleneck and allows the orchestrator to overcome previous failures during deployment.

A framework where a robot orchestrator and its policy co-evolve during deployment by curating demonstrations to fine-tune the policy for recurring failures.

loop fit 9/10Xilun Zhang, Maggie Wang, Erik Bauer et al.via Hugging Face Papers
Co-Evolving Robot Orchestrators and Policies through Deployment
ProjectGitHub·24 Jun 2026·Tools and harness

Ker102/Harneloop

The loop: Self-improving AI agents use trace-backed diagnosis and artifact-aware testing to refine their own harness units. This creates a better execution environment for their subsequent runs.

An open-source framework for building self-improving AI agent harnesses with portable units and evidence-gated promotion.

loop fit 7/10Ker102via GitHub
Ker102/Harneloop
ProjectGitHub·6 Oct 2026·Research automation

richardcsuwandi/kernaut

The loop: An automated research agent searches for novel GPU kernels. These kernels can then be used to discover and train more efficient architectures for future versions of the agent.

A repository for automated research of GPU kernels to enable open-ended model discovery.

loop fit 6/10richardcsuwandivia GitHub
richardcsuwandi/kernaut
Forum postLessWrong·8 Oct 2026·Interpretability and oversight

J++ Lens: Jacobian Filtering Enables More Faithful Workspace Lenses

The loop: The J++ Lens provides more faithful readouts of model activations, which improves the ability to monitor and oversee the training of future models.

Kola Ayonrinde, Anthropic Fellows, koayon@gmail.com; Jack Lindsey, Anthropic TL;DR: The J-Lens was proposed to read verbalisable representations from language model activations. We introduce the J++ Lens: an improvement to the J-Lens that...

loop fit 7/10Kola Ayonrindevia LessWrong
J++ Lens: Jacobian Filtering Enables More Faithful Workspace Lenses
Blog postAnthropic·8 Oct 2026·Tools and harness

An opt-in vulnerability-finding service for open-source software

The loop: Claude scans open source software for vulnerabilities, which improves the security of the libraries Claude uses, making the system more robust.

We’re launching OSS Scanner, an opt-in vulnerability scanner for the open-source ecosystem informed by our experience using Claude to find vulnerabilities during Project Glasswing. Projects that join will receive thorough, periodic...

loop fit 6/10via Anthropic
An opt-in vulnerability-finding service for open-source software

← 8 Oct 2026 Latest