Alphabell.
Foundation · Architecture and optimizer search

NAS with RL

Neural Architecture Search with Reinforcement LearningA recurrent network controller generates descriptions of neural network architectures and is trained with reinforcement learning to maximise the validation accuracy of the networks it designs. It produced a CIFAR-10 image model and a Penn Treebank recurrent cell that matched or beat human-designed ones.

The loop

An RNN controller writes the architecture of a child neural network; the child is trained and its validation accuracy is returned to the controller as a reward, and policy-gradient updates make the controller propose better architectures in the next round. The output is a better AI model designed by an AI model; the paper does not use the discovered architectures inside the controller itself, which is the remaining step to recursion.

The loop
NAS with RL
Foundation · arXiv
  1. RNN controller writes child network architecture
  2. Child validation accuracy returns as reward
  3. Updates make controller propose better architectures
  1. RNN controller writes child network architecture
  2. Child validation accuracy returns as reward
  3. Updates make controller propose better architectures
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

It established that a neural network can be trained to design other neural networks, so the designer can in principle be rebuilt from the architectures it discovers.

Evidence

Starting from scratch, the controller designed a CIFAR-10 network with a 3.65 test error rate (0.09 percent better and 1.05x faster than the previous comparable state of the art) and a Penn Treebank cell with 62.4 test perplexity, 3.6 better than the prior state of the art.

neural-architecture-searchreinforcement-learningautomlgoogle-brain