Alphabell.
Daily edition

Radar, 5 Oct 2026

Today's edition highlights diverse approaches to self-improvement, from multimodal models verifying their own outputs to optimizers evolving their own diagnostic tools.

PaperarXiv·2 Oct 2026·Training data

Recursive Self-Improvement in Unified Multimodal Models

The loop: A unified multimodal model generates images and writes programs to evaluate them, using the verified results to train its own visual understanding and generation.

A unified multimodal model uses program execution to verify its own generated images, creating a reliable training loop that improves its visual and text capabilities over multiple rounds.

loop fit 10/10Huijuan Wang, Chufan Shi, Cheng Yang et al.via arXiv
Recursive Self-Improvement in Unified Multimodal Models
PaperarXiv·2 Oct 2026·Tools and harness

VERSE: Verified Self-Evolving Optimizer for Agent Harnesses

The loop: The VERSE optimizer tests draft edits and replays failures to revise an agent harness alongside its own prompts and tools, improving its ability to diagnose and fix future errors.

An agent optimizer improves its own diagnostic tools and workflows while evolving an executor's harness, using execution-based verification to prevent regressions.

loop fit 9/10Zekai Wang, Yingqiang Ge, Zekun Wang et al.via arXiv
VERSE: Verified Self-Evolving Optimizer for Agent Harnesses
PaperarXiv·2 Oct 2026·Training data

Recursive Harness Self-Improvement for Frontier Reasoning Data Synthesis

The loop: A data synthesis system recursively improves its own generation harness by converting intermediate solver failures into reusable skills, producing progressively harder reasoning tasks.

A recursive self-improvement framework co-evolves its reasoning-data synthesis harness alongside the tasks it generates, significantly increasing task difficulty and downstream model performance.

loop fit 9/10Wenlong Zhang, Zhengbo Jiao, Chenxu Zhang et al.via arXiv
Recursive Harness Self-Improvement for Frontier Reasoning Data Synthesis
PaperarXiv·29 Sep 2026·Self-reward and self-play

AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation

The loop: An advisor model reflects on completed interactions to propose corrections, then uses targeted self-distillation to improve the advice it issues to a frozen language-model executor.

A small trainable advisor improves its ability to steer a frozen language model by selectively self-distilling from its own reflection on past interactions.

loop fit 9/10Rishabh Agrawal, Hejie Cui, Shasha Li et al.via Hugging Face Papers
AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation
Forum postLessWrong·4 Oct 2026·Training data

Reducing Synthetic Markers Makes Some SDF False Facts Linearly Indistinguishable from Pretraining-Acquired Knowledge

The loop: A synthetic document generator produces training data with reduced markers, which finetunes a model to internalize false facts that evade middle-layer linear probes.

Researchers demonstrate that reducing synthetic markers in generated training documents allows finetuning to implant false facts that are indistinguishable from pretraining-acquired knowledge.

loop fit 6/10Jason Zengvia LessWrong
Reducing Synthetic Markers Makes Some SDF False Facts Linearly Indistinguishable from Pretraining-Acquired Knowledge
Blog postPyTorch·1 Oct 2026·Kernels and compilers

Optimizing Jagged Flash Attention with TLX: The Road Toward SOTA FA4 on Blackwell

The loop: The Jagged Flash Attention kernel improves inference efficiency on Blackwell, which makes the Generative Ads Model faster.

TL;DR In this blog post, we present our work on Jagged Flash Attention (JFA) - the attention kernel behind Meta’s Generative Ads Model (GEM) - on NVIDIA Blackwell (B200), built...

loop fit 6/10Han Xu, Jacky Zhou, Jackie (Jiaqi) Xu, Hongtao Yu, Peng Chen (Dev Infra),...via PyTorch
Optimizing Jagged Flash Attention with TLX: The Road Toward SOTA FA4 on Blackwell
Blog postAi2·1 Oct 2026·Tools and harness

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

The loop: Olmo-core 3 provides an open training stack that improves the efficiency of training MoE models, which are then used to build better versions of the stack.

Olmo-core 3 introduces a redesigned, fully open training stack for efficiently scaling mixture-of-experts models into the trillion-parameter range.

loop fit 6/10via Ai2
Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

← 4 Oct 2026 Latest