Alphabell.
Daily edition

Radar, 10 Oct 2026

Today's radar highlights co-evolutionary loops, from web agents hardening against prompt injections to the joint optimization of data synthesizers and reasoning models.

PaperarXiv·7 Oct 2026·Training data

AdvSim2Real: Training Web Agents Against Adaptive Prompt Injection in a Web World Model

The loop: A web world model co-evolves a task curriculum and an injection adversary to train a web agent. The resulting training data improves the agent's robustness and capability, allowing it to handle more difficult tasks and stronger adversaries in the next iteration.

A training framework co-evolves a task curriculum and an injection adversary inside a web world model to make a web agent more capable and robust against prompt injections.

loop fit 10/10Sarim Hashmi, Mukul Ranjan, Kshitij Mishra, Mikhail Kuznetsov, Praneeth...via arXiv
AdvSim2Real: Training Web Agents Against Adaptive Prompt Injection in a Web World Model
PaperarXiv·6 Oct 2026·Architecture and optimizer search

RLDiscover: LLM-Driven Co-Evolution of Reinforcement Learning Algorithms

The loop: The RLDiscover framework progressively co-evolves the components of model-free deep reinforcement learning algorithms. The discovered algorithms improve the learning efficiency of the agents that use them, providing better fitness signals for the framework to discover even stronger algorithms.

An LLM-driven framework progressively co-evolves the components of deep reinforcement learning algorithms to discover variants that substantially improve agent learning efficiency.

loop fit 9/10Haoran Li, Zengle Ge, Xiaomin Yuan et al.via arXiv
RLDiscover: LLM-Driven Co-Evolution of Reinforcement Learning Algorithms
PaperarXiv·8 Oct 2026·Tools and harness

Harness Evolution Hits a Ceiling: When Weight Training Should Begin

The loop: A self-evolving harness loop repairs process failures to generate successful execution trajectories for an LLM agent. These trajectories are then used to train the model's weights, internalizing the gains and making the model better under the original harness.

A study shows that self-evolving agent harnesses excel at fixing process failures, generating successful trajectories that can then be trained into the model's weights to address content failures.

loop fit 9/10Yuan Tian, Bing Hu, Hao Wang et al.via arXiv
Harness Evolution Hits a Ceiling: When Weight Training Should Begin
PaperarXiv·6 Oct 2026·Self-modifying agents

Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents

The loop: A self-improving agent constructs meta-experience by re-executing incumbent and revised meta-skills from the same restored discovery state. This hindsight distillation improves the agent's meta-skills, making it better at discovering and refining future task-skills.

A mechanism for self-improving agents isolates the true impact of meta-skill revisions by re-executing them from identical states, improving how the agent discovers future skills.

loop fit 9/10Qianhan Feng, Zhongzhen Huang, Yakun Zhu et al.via arXiv
Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents
PaperarXiv·8 Oct 2026·Training data

SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning

The loop: A Synthesizer agent constructs training tasks based on a Reasoner agent's current capabilities, and the Reasoner learns from the resulting experience. The outcomes of the Reasoner's rollouts provide complementary rewards that jointly optimize both agents, making the Synthesizer better at generating informative tasks.

A multi-agent reinforcement learning framework jointly optimizes a synthesizer that generates training tasks and a reasoner that learns from them, ensuring tasks remain informative as the reasoner evolves.

loop fit 9/10Wei Yang, Shawn Li, Yuehan Qin et al.via arXiv
SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning

← 9 Oct 2026 Latest