Alphabell.
Data and training signal

Self-reward and self-play

Models that supply their own training signal: judges, critics, verifiers and self-play opponents.

Recently on the radar

More →
PaperarXiv·29 Sep 2026·Self-reward and self-play

RLTL;DR: Self-improvement by Internalizing Self-generated Feedback

The loop: A policy generates its own feedback insights from failed attempts at difficult tasks. It then internalizes these insights through training, breaking through learning barriers and improving its success rate on future rollouts.

loop fit 9/10Michael Kirchhof, Eleonora Gualdoni, Andrew Szot et al.via arXiv

Other loop types