Alphabell.
Updated daily

Daily radar

New work where AI improves AI, picked each day from public sources by our editor pipeline. Every item states its loop and links to the original.

What today's radar picked, by part of the AI stack
Agents and harness 2Compute stack 2Data and training signal 4

Last run 2 Oct 2026, 18:24 UTC · 31 of 31 sources checked · 160 candidates scored · 8 published · RSS · Sources · How it works

2 Oct 2026

8 picked
PaperarXiv·29 Sep 2026·Inference efficiency

Context Language Models

The loop: A language model natively manages its own context by updating a context file. The model's context-management instructions are then evolved through a skill-optimization loop, improving its accuracy and efficiency on subsequent tasks.

Language models are designed to natively manage their own context as a file, enabling intrinsic context management that can be optimized through skill evolution.

loop fit 10/10Ru-Lin Shao, Shannon Zejiang Shen, J. Yin et al.via Semantic Scholar
PaperarXiv·28 Sep 2026·Kernels and compilers

AReaL-TIK: Stateful Agentic Optimization of Unified RL Kernels through an Optimization IR

The loop: The KernelBraid agent optimizes unified RL kernels through an intermediate representation and verifies the improvements. These faster kernels are then used to accelerate the reinforcement learning training process.

An agentic framework optimizes unified RL kernels for rollout and policy update, preserving bitwise consistency and improving aggregate latency.

loop fit 9/10Ran Yan, Youhe Jiang, Jiayi Nie et al.via arXiv
PaperarXiv·29 Sep 2026·Self-reward and self-play

Train Ahead, Distill Back: Bootstrapping On-Policy Self-Distillation for Large Language Models

The loop: A policy is temporarily trained ahead to create a stronger future teacher. This future teacher then provides dense token-level supervision to the restarted original policy, improving its performance.

A self-distillation method temporarily trains a model ahead to create a future teacher that provides better supervision for the original model.

loop fit 9/10Zheng Zhang, Xinyue Tan, Lufei Li et al.via arXiv
PaperarXiv·29 Sep 2026·Tools and harness

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

The loop: Mutator agents rewrite complete agent harnesses using global search history and feedback. An orchestrator adapts the search strategy based on the performance of the discovered harnesses, improving the mutators' future harness designs.

A framework co-evolves agent harnesses and the mutator agents that design them, using hierarchical memory and orchestrated search adaptation.

loop fit 9/10Prithwish Jana, Mononito Goswami, Hao Liu et al.via arXiv
PaperarXiv·30 Sep 2026·Training data

False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

The loop: Self-evolving search agents generate their own training data, which often suffers from co-cheating where the proposer and solver agree on shared errors. This data is filtered using multi-sample verification and cross-fitting to reliably train the next generation of agents.

A study identifies and mitigates co-cheating in self-evolving search agents by verifying and partitioning self-generated training data.

loop fit 9/10Meijia Chen, Hao Li, Zheng Lu et al.via arXiv
PaperarXiv·29 Sep 2026·Self-modifying agents

SelfSearch: Reward-Free Search for Self-Improving Agents

The loop: An agent modifies its own instructions and tools using records of its previous self-improvement episodes. These modifications directly improve the agent's success rate on future downstream tasks without requiring external reward signals during the search.

A reward-free search procedure allows agents to modify their own instructions and tools using records of past self-improvement attempts.

loop fit 9/10Jungwoo Yang, Injin Kong, Yohan Jovia arXiv
PaperarXiv·29 Sep 2026·Self-reward and self-play

RLTL;DR: Self-improvement by Internalizing Self-generated Feedback

The loop: A policy generates its own feedback insights from failed attempts at difficult tasks. It then internalizes these insights through training, breaking through learning barriers and improving its success rate on future rollouts.

A reinforcement learning approach allows a policy to generate and internalize its own feedback insights from failed attempts to improve performance on difficult tasks.

loop fit 9/10Michael Kirchhof, Eleonora Gualdoni, Andrew Szot et al.via arXiv
PaperarXiv·1 Oct 2026·Training data

No Model Required: Text Entropy Rate Filtering Mitigates Iterative Fine-Tuning Collapse

The loop: A model generates synthetic text which is then filtered using a non-parametric entropy rate estimator to preserve diversity. This filtered data is used to iteratively fine-tune the model, preventing model collapse across successive generations.

A non-parametric entropy rate estimator filters synthetic text to maintain diversity and prevent model collapse during iterative fine-tuning.

loop fit 9/10Lewis Mitchellvia arXiv