Alphabell.
Paper · Training data

SEAL

Self-Adapting Language Models (SEAL)SEAL is a framework in which a language model writes self-edits, its own finetuning data and update directives, and applies them to its own weights through supervised finetuning.

The loop

The language model generates a self-edit for a new input (restructured training data, optimization hyperparameters or augmentation calls), and the edit is applied as a persistent weight update through SFT. Reinforcement learning then rewards each self-edit by the downstream performance of the updated model, so the model learns to write better training data for itself.

The loop
SEAL
Paper · arXiv
  1. Language model generates a self-edit
  2. Edit is applied as weight update
  3. RL rewards self-edit by downstream performance
  1. Language model generates a self-edit
  2. Edit is applied as weight update
  3. RL rewards self-edit by downstream performance
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

The model is both the author of its training data and the trainee, and the author is optimized by how much its own update helped, which is a direct self-modification loop over weights.

Evidence

In single-passage knowledge incorporation, two rounds of ReST-EM raised QA accuracy from 32.7% (no adaptation) to 47.0%, ahead of finetuning on raw passages or on synthetic data generated by GPT-4.1.

self-editingsynthetic-datacontinual-learningreinforcement-learningweight-updates