Self-training learned optimizers
Training Learned Optimizers with Randomly Initialized Learned OptimizersA population of randomly initialized learned optimizers (small neural networks that compute parameter updates) is used to meta-train the learned optimizers themselves, with no hand-designed optimizer such as Adam anywhere in the process. Population based training decides which optimizers survive and which train which.
Learned optimizers from the population act as the outer optimizer that updates the parameters of other learned optimizers, and population based training keeps the best-performing optimizer parameters together with the optimizer that trained them. As the optimizers get better at optimizing, they train each other faster, which the authors describe as a positive feedback loop.
- Learned optimizers update other learned optimizers
- Training keeps best performing optimizer parameters
- Better optimizers train each other faster
- Learned optimizers update other learned optimizers
- Training keeps best performing optimizer parameters
- Better optimizers train each other faster
Why it is a road to recursion
The improved artifact (the optimizer) is the same tool that trains the next improved artifact, a literal self-referential loop in which gains compound.
Evidence
After 10 days on a 500K CPU core cluster, the self-trained population went from diverging or learning slowly to producing learned optimizers that outperform per-problem learning-rate-tuned Adam on the training distribution of tasks.
Related loops
More architecture and optimizer search →- LLM ensemble rewrites candidate programs
- Evaluator scores and selects best programs
- Evolved mixture-of-experts load-balancing loss
- Loss improves LLM training
- repeat
- Designer agents propose language model architectures
- Verifier agents pre-train and evaluate designs
- Verified results feed evolutionary population
- Language models search for better architectures
- repeat
- GPT-4 proposes preference-optimization losses
- Candidate losses fine-tune language models
- Evaluation scores feed back into prompt
- Best loss aligns other LLMs
- repeat