Why adding the input back lets networks go deep
A residual connection is a wire that carries a layer's input past it and adds it back to the output. The layer no longer has to produce the answer, only the correction. That one change took networks from a few dozen layers to more than a hundred.
Deeper got worse
By 2015 the best image classifiers had sixteen to thirty layers; would more keep helping? He, Zhang, Ren and Sun tried it and found the opposite. On ImageNet, their 34-layer plain network — one layer after another, nothing else — had higher training error than their 18-layer plain network, all through training. That is strange: a deeper network can, in principle, do anything the shallower one does — set the extra layers to pass their input through unchanged and copy the rest. So it is not short of capacity, nor memorising, since it is worse on the training set itself. The authors call this the degradation problem and read it as one of optimisation: the solution exists, but training cannot find it.
Why it matters
Depth is where a network's power comes from — each layer builds on the one below — so a wall at thirty layers walled the whole field. In the same paper, residual networks up to 152 layers deep were trained on ImageNet; an ensemble of six residual nets of different depths scored 3.57% top-5 error on the test set and won the 2015 ImageNet classification contest.
Interactive Set the depth and the weight size, then switch skip connections on and off — the same random weights are used both ways, and the readout tells you how the size of the signal changes through the stack.
Learning the difference instead of the thing
Call the function a block should compute H(x). A plain block builds all of H from scratch. A residual block builds only F(x) = H(x) − x, the part that differs from the input, and the wire adds x back: output = F(x) + x. The wire has no parameters and costs one addition: a plain net and its residual twin share depth, width and cost. Only the target changes.
Take a block whose best behaviour is to do nothing. A plain block must learn the identity: several nonlinear layers cooperating to reproduce their input exactly, awkward for weights that start small and random. A residual block gets it free: zero weights, F vanishes, the output is x. The authors' hypothesis: the best block is usually close to the identity, and a small adjustment to a known reference is easier to find than a whole function from nothing. Their measurements agree: learned residual functions have generally small outputs, and the deeper the network, the less each layer changes the signal.
One caution the paper is careful about: its plain networks used batch normalisation, and the authors checked that neither forward signals nor backward gradients were vanishing. They do not attribute the degradation to fading gradients, leaving its cause as future work. The toy above is about the identity argument, not a model of their ImageNet runs.
What the numbers showed
With shortcuts added and nothing else changed, this reversed. The 34-layer residual network beat the 18-layer one by 2.8% with lower training error, and cut the 34-layer plain net's top-1 error by 3.5%. At 18 layers plain and residual were comparably accurate, the residual one just converging faster — which fits: shallow plain nets were never the problem. On CIFAR-10 the authors pushed to 1202 layers and reached training error under 0.1%, though its test error, 7.93%, was worse than their 110-layer network's, which they attribute to overfitting. Optimisation had stopped being the limit.
In short
More plain layers made networks harder to train, not merely no better. A residual connection adds each block's input back to its output, so the block only learns the difference from the identity; doing nothing becomes the easiest thing to learn. That is why the same layers, rewired, go a hundred and more deep.
Where this comes from
- Deep Residual Learning for Image Recognition linked only, not reproduced
arxiv.org/abs/1512.03385