Alphabell.
Daily edition

Radar, 4 Oct 2026

Today's edition highlights diverse approaches to self-improvement, from agents managing their own context windows to reflecting on internal representations for skill evolution.

PaperarXiv·27 Sep 2026·Self-modifying agents

Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents

The loop: The Rep2Skill agent models its own internal representation trajectories to localize execution errors, generating textual feedback to evolve its external skills.

A representation-guided framework allows LLM agents to improve their external textual skills by reflecting on their own internal execution state representations.

loop fit 9/10Euntae Choi, Su-Min Song, Sungjoo Yoovia Semantic Scholar
Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents
PaperarXiv·29 Sep 2026·Inference efficiency

AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

The loop: A coding agent is trained to decide when and how to compact its own context during long-horizon tasks, using task-success rewards to improve its policy.

A framework trains long-horizon coding agents to manage their own context by deciding when to compact and what working state to preserve, optimizing these decisions through reinforcement learning.

loop fit 9/10Jitin Singla, Parikshit Pareek, Pratik Jawanpuria et al.via arXiv
AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents
PaperarXiv·1 Oct 2026·Tools and harness

LiteEvo: Automated, Cost-Efficient Harness Evolution for Generalization to Unseen Tasks

The loop: LiteEvo meta-agents mine agent trajectories for reusable components to evolve a harness library, which improves the agent's performance on unseen tasks.

A lightweight harness-evolution algorithm uses tool-free meta-agents to mine trajectories and build a reusable component library, improving agent performance on unseen tasks at lower cost.

loop fit 9/10Geyi Yang, Zikun Qu, Xiang Li et al.via arXiv
LiteEvo: Automated, Cost-Efficient Harness Evolution for Generalization to Unseen Tasks
PaperarXiv·30 Sep 2026·Self-reward and self-play

AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation

The loop: An advisor model uses reflection to propose corrections to its own decisions, then uses self-distillation from a feedback-conditioned copy to improve its future advice.

A method for steering frozen language models pairs outcome-based reinforcement learning with selective self-distillation to improve a trainable advisor's natural-language advice.

loop fit 9/10Zhijie Wei, Ferris Tan, Jinghui Wangvia arXiv
AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation
Blog postNVIDIA·28 Sep 2026·Hardware and chips

How NVIDIA DSX MaxLPS Maximizes AI Factory Throughput and Efficiency

The loop: DSX MaxLPS optimizes power usage in AI factories, which improves throughput for the models running in that factory.

Every unused watt is capacity left on the table. AI factories are typically provisioned for the unlikely moment when every GPU reaches peak power, creating a...

loop fit 7/10Sarah McKenneyvia NVIDIA Technical Blog
How NVIDIA DSX MaxLPS Maximizes AI Factory Throughput and Efficiency
Blog postPyTorch·30 Sep 2026·Tools and harness

From Upstream Changes to Downstream Confidence: Inside Torch Spyre’s Integration with PyTorch CRCR

The loop: Torch Spyre uses the CRCR relay to automate CI testing, which improves the stability of PyTorch for its own development.

TL;DR PyTorch’s Cross-Repository CI Relay (CRCR) gives out-of-tree accelerators a clean, scalable way to plug into upstream CI - and it leaves each backend free to decide which of PyTorch’s...

loop fit 7/10Mehant Kammakomati (IBM), Jewel K M (Red Hat), Anubhav Jana (IBM), Padmanabha...via PyTorch
From Upstream Changes to Downstream Confidence: Inside Torch Spyre’s Integration with PyTorch CRCR
Blog postOpenAI·28 Sep 2026·Interpretability and oversight

Towards safety cases for frontier AI training

The loop: A safety case framework investigates misalignment incidents, which improves technical safeguards for the training of successor frontier models.

Our early guidelines for safety cases in frontier AI training cover technical safeguards, operational practices, and investigating misalignment incidents

loop fit 6/10via OpenAI
Towards safety cases for frontier AI training

← 3 Oct 2026 Latest