Alphabell.
Daily edition

Radar, 11 Oct 2026

Today's edition highlights the growing trend of co-evolution, with multiple projects demonstrating how models can simultaneously improve their weights alongside their harnesses, rubrics, or test-time policies.

PaperarXiv·7 Oct 2026·Self-reward and self-play

Recurrent Self-Improvement: Dynamic Cross-Loop On-Policy Distillation for Looped Language Models

The loop: A Looped Language Model uses its own extended recurrent computation as a teacher to supervise its intermediate loops. Because the parameters are shared across all loops, distilling into the intermediate steps updates the shared weights, which immediately improves the terminal loop teacher for the next round of self-improvement.

A cross-loop distillation framework where a looped language model uses its own deeper recurrent computations to supervise and improve its shared parameters.

loop fit 9/10Yi Wang, Rui Qian, Yu Li et al.via arXiv
Recurrent Self-Improvement: Dynamic Cross-Loop On-Policy Distillation for Looped Language Models
PaperarXiv·8 Oct 2026·Self-modifying agents

Agentic-TTT: Training Test-Time Policy for Test-Time Training

The loop: A language model learns a test-time policy to decide when and how to apply test-time training algorithms to itself. By training this policy on the utility gains observed from its past decisions, the model becomes better at managing its own parameter-level self-improvement during deployment.

A framework that gives a model the agency to select and apply test-time training algorithms to itself based on observed utility gains.

loop fit 9/10Jiahao Lu, Mohan Kankanhallivia arXiv
Agentic-TTT: Training Test-Time Policy for Test-Time Training
PaperarXiv·7 Oct 2026·Tools and harness

CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution

The loop: A terminal agent framework alternates between using execution failures to synthesize a better runtime harness and using verified rollouts to train the model policy. This co-evolution ensures that the model is trained on data matched to its adopted runtime, improving both the agent's weights and the harness it depends on.

An alternating co-evolution framework that decouples harness search and policy training to systematically improve both components of a terminal agent.

loop fit 9/10Jixuan Chen, Jiaxin Zhang, Qinyuan Ye et al.via arXiv
CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution
PaperarXiv·9 Oct 2026·Self-reward and self-play

EvoRubric: Self-Evolving Rubric-Driven RL for Open-Ended Generation

The loop: A shared language model policy acts as both a reasoner generating responses and a generator creating evaluation rubrics. Discriminative feedback and peer consensus are used to improve the rubrics, which in turn provide better complementary rewards to optimize the shared policy for open-ended tasks.

A co-evolutionary reinforcement learning framework where a model iteratively discovers and refines the evaluation rubrics used to train its own policy.

loop fit 9/10Xin Guan, Xiaomeng Hu, Shen Huang, Zhenyi Wang, Bo Zhang, Zijian Li, Pengjun...via arXiv
EvoRubric: Self-Evolving Rubric-Driven RL for Open-Ended Generation
PaperarXiv·9 Oct 2026·Tools and harness

DuplexAgent-RSI: Recursive Harness Improvement for Full-Duplex Voice Agent Collaboration

The loop: A full-duplex voice agent system uses interaction traces from simulated conversations to systematically revise its own harness modules. These revisions improve how the system coordinates task delegation and responsiveness, making the agent better at handling complex live collaborations.

A closed-loop system that uses simulated interaction traces to recursively improve the coordination harness of a full-duplex voice agent.

loop fit 9/10Yingda Shen, Yuxiang Wang, Kunyu Feng, Qinke Ni, Jiaqi Li, Minghao Hsu, Junan...via arXiv
DuplexAgent-RSI: Recursive Harness Improvement for Full-Duplex Voice Agent Collaboration
PaperarXiv·9 Oct 2026·Architecture and optimizer search

Neural Architecture Discovery via Autonomous Evolution

The loop: The ASI-Arch system autonomously conducts neural architecture research through a closed loop process of experimenting, analyzing, and updating. This discovers novel architectures that can be used to train more capable base models, which in turn power future versions of the research agent.

An autonomous research agent that iteratively designs, tests, and refines neural architectures, discovering variants that significantly outperform existing models.

loop fit 8/10Weixian Xu, Yixiu Liu, Yang Nan, Lyumanshan Ye, Xiangkun Hu, Zhen Qin, Pengfei...via arXiv
Neural Architecture Discovery via Autonomous Evolution

← 10 Oct 2026 Latest