STOP
Self-Taught Optimizer (STOP): Recursively Self-Improving Code GenerationA seed 'improver' program, which calls a language model to improve a given program, is run on its own source code to produce better improvers. Only the scaffolding changes; the language model weights are not altered.
GPT-4, called from inside the improver program, rewrites that same improver and each rewrite is scored by a meta-utility: how well the new improver optimizes code on downstream tasks. The best rewrite becomes the improver for the next round, and the model proposed strategies such as beam search, genetic algorithms and simulated annealing on its own.
- GPT-4 rewrites the improver program
- Each rewrite is scored by meta-utility
- Best rewrite becomes next round improver
- GPT-4 rewrites the improver program
- Each rewrite is scored by meta-utility
- Best rewrite becomes next round improver
Why it is a road to recursion
It demonstrates the minimal recursive step, a program that improves programs improving itself, and the authors state that full recursion would also require improving the language model.
Evidence
An improver self-improved for four rounds on learning parity with noise transferred to unseen tasks, for example raising the 3-SAT score from 21.2% with the seed improver to 75.1%.
Related loops
More self-modifying agents →- Coding agent edits its own code
- Scored child agents enter the archive
- Archived agents make later self modifications
- repeat
- Claude Code writes its own code changes
- Changes become next Claude Code versions
- Release becomes harness for next development round
- repeat
- Outer agent rewrites inner agent code
- Accepted rewrites become the incumbent agent
- Discovered agent becomes the outer loop agent
- repeat