№ 45 · machine learning

Why switching neurons off makes a network learn better

For each training example, flip a coin for every hidden neuron and switch off the losers. The network that learns this way is worse at every single step — and better at the end.

Switching neurons off on purpose

A neural network is layers of small units, each passing on a weighted sum of the layer below. Dropout is a training rule: on each training case, every hidden unit is removed at random with some probability — 0.5 in the original paper. The rest do the forward pass, the weights update as usual, and the next case removes a fresh random set. At test time nothing is removed.

Why it matters

The wall is overfitting: a large network trained on modest data can fit the training set almost perfectly in many ways, most of which do worse on new data. Hinton and colleagues report that on the MNIST handwritten digits, the best published plain feedforward network made 160 test errors; the same network with 50% dropout of the hidden units — plus a separate cap on the size of each hidden unit's incoming weights — made about 130, and about 110 when a random 20% of the input pixels were dropped too. The rule change is tiny; the cost is longer training — their ImageNet net took roughly twice as long with dropout, about four days against two.

Interactive Press train step to flip the coins and see one thinned network; change drop probability and hidden units and watch the count of possible thinned networks, the test-time scaling, and how close the mean network gets to the average of all of them.

An illustrative network with fixed, arbitrary weights and one fixed input — it is not trained, so no error curve is shown. Each train step drops every hidden unit independently with the chosen probability and reads the output of the thinned network that remains. All 2n thinned networks is computed by enumerating every mask and averaging their outputs. The mean network keeps every unit and multiplies its outgoing weights by 1 − p (a half at p = 0.5, the paper's convention); with a sigmoid output this equals the normalised geometric mean of the thinned networks' predictions, which is the exact result the paper proves for one hidden layer and a softmax, and it sits close to — but not exactly on — their plain arithmetic mean. The inverted-dropout factor 1/(1 − p) is the equivalent correction applied during training instead — a later convention, not one used in this paper, which only drops at 0.5 and halves at test time. The geometric mean shown is computed directly: the probability-weighted product of each thinned network's two class probabilities, then normalised. The running average over your sampled steps drifts toward the exact average as you take more of them.

Two ways to see what the coin does

The first view is about partnerships. Without dropout, a hidden unit can learn to be useful only alongside a few specific other units, producing a signal meaningless on its own that its partners correct. The paper calls this co-adaptation. Such partnerships rarely survive new data. With dropout, half the partners are missing on any given case, so a unit is only useful if it detects something that helps in many contexts. The authors report that features learned under dropout look like simple pen strokes; those learned without it are hard to interpret.

The second view is about averaging. Training many networks and averaging their predictions reduces test error, but is expensive. Dropout gets something related almost for free. With N hidden units in a layer, there are 2N ways to thin it, so each training case trains a different thinned network, all sharing weights for whichever units are present. One set of weights trains an exponential number of networks at once.

The mean network

Rather than sample thinned networks at test time and average their answers, the paper proposes a shortcut: use the whole network and halve each unit's outgoing weights, since twice as many units are active as during training. This is the mean network. For one hidden layer and a softmax output, they show it is exactly the geometric mean of the predictions of all 2N thinned networks; for deeper networks it is an approximation that works about as well in practice.

That halving is specific to a drop probability of 0.5. If a fraction p is dropped, the same bookkeeping says to scale by 1 − p at test time. Many implementations today do it the other way round, dividing the surviving units by 1 − p during training so the test network needs no adjustment. The expected signal is the same either way; only the moment of scaling differs.

In short

Dropout trains a network while randomly deleting hidden units. No unit can rely on specific partners, so each learns something useful on its own. Each case trains a different thinned network from an exponentially large family sharing one set of weights. At test time, the full network with scaled weights stands in for that family's average.

Where this comes from

  1. Improving neural networks by preventing co-adaptation of feature detectors linked only, not reproduced
    Geoffrey E. Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, Ruslan R. Salakhutdinov · arXiv:1207.0580 · 2012
    arxiv.org/abs/1207.0580