Unit content
Residual connections and deep networks
As neural networks become deeper, training can become harder even when the additional layers increase representational capacity. Residual connections give information and gradients a short path around a block of layers.
A residual block has the form
$$\mathbf y=\mathbf x+F(\mathbf x).$$
Instead of learning the whole mapping directly, the block learns a residual change $F(\mathbf x)$ relative to its input.
Identity path
If the residual branch initially contributes little, the block behaves approximately like the identity function. Adding more blocks therefore need not immediately destroy an already useful representation.
Gradient flow
The additive shortcut provides a direct derivative path through the network. This helps useful gradient information reach earlier layers during training.
Shape matching
The tensors being added must have compatible shapes. When channel count or spatial resolution changes, a learned projection such as a $1\times1$ convolution can transform the shortcut before addition.
Residual networks
A ResNet stacks many residual blocks. Residual connections do not eliminate all optimization difficulties, but they made very deep convolutional networks substantially easier to train and have since become common far beyond computer vision.