Alphabell.
Daily edition

Radar, 6 Oct 2026

Today's edition highlights systems that optimize their own physical and software infrastructure, from agents rewriting their inference engines to language models designing sparse accelerators.

PaperarXiv·3 Oct 2026·Inference efficiency

SEIS: Self-Evolving Inference Systems

The loop: SEIS autonomously optimizes its own inference engine code through iterative self-evolution. This redesigns the engine for higher throughput, which accelerates the models that power the system.

An agentic system autonomously optimizes the mini-sglang inference engine end-to-end through iterative code changes and inherited experiences, achieving a 3.27X throughput speedup.

loop fit 10/10Zhen Xu, Jingyu Liu, Zongze Li et al.via arXiv
SEIS: Self-Evolving Inference Systems
PaperarXiv·6 Oct 2026·Hardware and chips

SparseCraft: Agentic Hardware-Software Co-Optimization for Sparse Computing

The loop: A language model iteratively edits the RTL and schedules of a sparse accelerator based on measured simulation outcomes. This produces a faster and more efficient hardware design, which can then run the model itself more efficiently.

A language model inside a closed loop iteratively edits the RTL, memory configuration, and sparse-kernel schedule of an accelerator, achieving significant cycle and area reductions.

loop fit 10/10Rajatabha Chakraborty, M P Samartha, Vedant Pahariya, Priyesh Shuklavia arXiv
SparseCraft: Agentic Hardware-Software Co-Optimization for Sparse Computing
PaperarXiv·4 Oct 2026·Tools and harness

MESH-Harness: Self-Improving Agent Harnesses via Bandit-Guided Compositional Evolution

The loop: An agent uses bandit-guided composition and local code edits to iteratively evolve its own harness modules. This produces a better harness that improves the agent's performance on future tasks.

An agent harness organizes itself into functional modules and uses bandit-guided composition and local code edits to iteratively evolve and improve its own configuration.

loop fit 9/10Zhiwei Shang, Yu Huo, Mingrong Gong et al.via arXiv
MESH-Harness: Self-Improving Agent Harnesses via Bandit-Guided Compositional Evolution
LeaderboardGitHub Pages·6 Oct 2026·AI R&D evals

New best on GSO: claude-fable-5-1_unknown · OpenHands (88.2%)

The loop: AI agents optimize real codebases for speed. This improves the software infrastructure that AI systems rely on, accelerating the execution of future agents.

Claude-fable-5.1 using OpenHands achieved a new state-of-the-art score of 88.2% on the GSO benchmark, which measures agents optimizing real codebases for speed.

New best on GSO: claude-fable-5-1_unknown · OpenHands (88.2%)
Forum postLessWrong·6 Oct 2026·Self-reward and self-play

Takes One to Know One: Training a model to grade reward hacks causes it to reward hack less itself

The loop: A grader model is fine-tuned to judge and catch reward hacks in other models. This training simultaneously improves the grader's own alignment, causing it to reward hack less itself.

Fine-tuning a model to act as a judge that catches reward hacks also reduces the model's own tendency to engage in reward hacking when completing tests.

loop fit 9/10Arjun Srivia LessWrong
Takes One to Know One: Training a model to grade reward hacks causes it to reward hack less itself
Blog postEpoch AI·2 Oct 2026·Research automation

Coding-agent use at OpenAI is doubling roughly every month

The loop: Coding agents assist OpenAI researchers in automating the development of AI software. This accelerates the engineering of the next generation of models and agents.

According to breakpoint fits of API spending data, the usage of coding agents by OpenAI researchers has recently been doubling roughly every month.

loop fit 6/10via Epoch AI
Coding-agent use at OpenAI is doubling roughly every month

← 5 Oct 2026 Latest