NAS with RL
Neural Architecture Search with Reinforcement LearningA recurrent network controller generates descriptions of neural network architectures and is trained with reinforcement learning to maximise the validation accuracy of the networks it designs. It produced a CIFAR-10 image model and a Penn Treebank recurrent cell that matched or beat human-designed ones.
An RNN controller writes the architecture of a child neural network; the child is trained and its validation accuracy is returned to the controller as a reward, and policy-gradient updates make the controller propose better architectures in the next round. The output is a better AI model designed by an AI model; the paper does not use the discovered architectures inside the controller itself, which is the remaining step to recursion.
- RNN controller writes child network architecture
- Child validation accuracy returns as reward
- Updates make controller propose better architectures
- RNN controller writes child network architecture
- Child validation accuracy returns as reward
- Updates make controller propose better architectures
Why it is a road to recursion
It established that a neural network can be trained to design other neural networks, so the designer can in principle be rebuilt from the architectures it discovers.
Evidence
Starting from scratch, the controller designed a CIFAR-10 network with a 3.65 test error rate (0.09 percent better and 1.05x faster than the previous comparable state of the art) and a Penn Treebank cell with 62.4 test perplexity, 3.6 better than the prior state of the art.
Related loops
More architecture and optimizer search →- Learned optimizers update other learned optimizers
- Training keeps best performing optimizer parameters
- Better optimizers train each other faster
- repeat
- LLM ensemble rewrites candidate programs
- Evaluator scores and selects best programs
- Evolved mixture-of-experts load-balancing loss
- Loss improves LLM training
- repeat
- Designer agents propose language model architectures
- Verifier agents pre-train and evaluate designs
- Verified results feed evolutionary population
- Language models search for better architectures
- repeat