Recurrent Self-Improvement: Dynamic Cross-Loop On-Policy Distillation for Looped Language Models
The loop: A Looped Language Model uses its own extended recurrent computation as a teacher to supervise its intermediate loops. Because the parameters are shared across all loops, distilling into the intermediate steps updates the shared weights, which immediately improves the terminal loop teacher for the next round of self-improvement.
A cross-loop distillation framework where a looped language model uses its own deeper recurrent computations to supervise and improve its shared parameters.





