·29 Sep 2026·Self-reward and self-play
Train Ahead, Distill Back: Bootstrapping On-Policy Self-Distillation for Large Language Models
The loop: A policy is temporarily trained ahead to create a stronger future teacher. This future teacher then provides dense token-level supervision to the restarted original policy, improving its performance.