Alphabell.
Paper · Self-modifying agents

STOP

Self-Taught Optimizer (STOP): Recursively Self-Improving Code GenerationA seed 'improver' program, which calls a language model to improve a given program, is run on its own source code to produce better improvers. Only the scaffolding changes; the language model weights are not altered.

The loop

GPT-4, called from inside the improver program, rewrites that same improver and each rewrite is scored by a meta-utility: how well the new improver optimizes code on downstream tasks. The best rewrite becomes the improver for the next round, and the model proposed strategies such as beam search, genetic algorithms and simulated annealing on its own.

The loop
STOP
Paper · arXiv
  1. GPT-4 rewrites the improver program
  2. Each rewrite is scored by meta-utility
  3. Best rewrite becomes next round improver
  1. GPT-4 rewrites the improver program
  2. Each rewrite is scored by meta-utility
  3. Best rewrite becomes next round improver
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

It demonstrates the minimal recursive step, a program that improves programs improving itself, and the authors state that full recursion would also require improving the language model.

Evidence

An improver self-improved for four rounds on learning parity with noise transferred to unseen tasks, for example raising the 3-SAT score from 21.2% with the seed improver to 75.1%.

scaffoldingself-improvementcode-generationgpt-4sandbox-safety