Why switching neurons off makes a network learn better
For each training example, flip a coin for every hidden neuron and switch off the losers. The network that learns this way is worse at every single step — and better at the end.
Switching neurons off on purpose
A neural network is layers of small units, each passing on a weighted sum of the layer below. Dropout is a training rule: on each training case, every hidden unit is removed at random with some probability — 0.5 in the original paper. The rest do the forward pass, the weights update as usual, and the next case removes a fresh random set. At test time nothing is removed.
Why it matters
The wall is overfitting: a large network trained on modest data can fit the training set almost perfectly in many ways, most of which do worse on new data. Hinton and colleagues report that on the MNIST handwritten digits, the best published plain feedforward network made 160 test errors; the same network with 50% dropout of the hidden units — plus a separate cap on the size of each hidden unit's incoming weights — made about 130, and about 110 when a random 20% of the input pixels were dropped too. The rule change is tiny; the cost is longer training — their ImageNet net took roughly twice as long with dropout, about four days against two.
Interactive Press train step to flip the coins and see one thinned network; change drop probability and hidden units and watch the count of possible thinned networks, the test-time scaling, and how close the mean network gets to the average of all of them.
Two ways to see what the coin does
The first view is about partnerships. Without dropout, a hidden unit can learn to be useful only alongside a few specific other units, producing a signal meaningless on its own that its partners correct. The paper calls this co-adaptation. Such partnerships rarely survive new data. With dropout, half the partners are missing on any given case, so a unit is only useful if it detects something that helps in many contexts. The authors report that features learned under dropout look like simple pen strokes; those learned without it are hard to interpret.
The second view is about averaging. Training many networks and averaging their predictions reduces test error, but is expensive. Dropout gets something related almost for free. With N hidden units in a layer, there are 2N ways to thin it, so each training case trains a different thinned network, all sharing weights for whichever units are present. One set of weights trains an exponential number of networks at once.
The mean network
Rather than sample thinned networks at test time and average their answers, the paper proposes a shortcut: use the whole network and halve each unit's outgoing weights, since twice as many units are active as during training. This is the mean network. For one hidden layer and a softmax output, they show it is exactly the geometric mean of the predictions of all 2N thinned networks; for deeper networks it is an approximation that works about as well in practice.
That halving is specific to a drop probability of 0.5. If a fraction p is dropped, the same bookkeeping says to scale by 1 − p at test time. Many implementations today do it the other way round, dividing the surviving units by 1 − p during training so the test network needs no adjustment. The expected signal is the same either way; only the moment of scaling differs.
In short
Dropout trains a network while randomly deleting hidden units. No unit can rely on specific partners, so each learns something useful on its own. Each case trains a different thinned network from an exponentially large family sharing one set of weights. At test time, the full network with scaled weights stands in for that family's average.
Where this comes from
- Improving neural networks by preventing co-adaptation of feature detectors linked only, not reproduced
arxiv.org/abs/1207.0580