Alphabell.
Foundation · Architecture and optimizer search

Self-training learned optimizers

Training Learned Optimizers with Randomly Initialized Learned OptimizersA population of randomly initialized learned optimizers (small neural networks that compute parameter updates) is used to meta-train the learned optimizers themselves, with no hand-designed optimizer such as Adam anywhere in the process. Population based training decides which optimizers survive and which train which.

The loop

Learned optimizers from the population act as the outer optimizer that updates the parameters of other learned optimizers, and population based training keeps the best-performing optimizer parameters together with the optimizer that trained them. As the optimizers get better at optimizing, they train each other faster, which the authors describe as a positive feedback loop.

The loop
Self-training learned optimizers
Foundation · arXiv
  1. Learned optimizers update other learned optimizers
  2. Training keeps best performing optimizer parameters
  3. Better optimizers train each other faster
  1. Learned optimizers update other learned optimizers
  2. Training keeps best performing optimizer parameters
  3. Better optimizers train each other faster
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

The improved artifact (the optimizer) is the same tool that trains the next improved artifact, a literal self-referential loop in which gains compound.

Evidence

After 10 days on a 500K CPU core cluster, the self-trained population went from diverging or learning slowly to producing learned optimizers that outperform per-problem learning-rate-tuned Adam on the training distribution of tasks.

learned-optimizersmeta-learningpopulation-based-traininggoogle-brain