Unit content
Convolutional neural networks
Images and other grid-like data contain local patterns that can appear at many positions. A convolutional neural network (CNN) uses shared local filters so the same feature detector can be applied across the input.
Convolutional filters
A small kernel slides across the input. At each position, the kernel weights are multiplied by nearby input values and summed to produce one output value.
For multi-channel data, each output channel combines information from all relevant input channels.
Weight sharing
The same kernel parameters are reused at every spatial position. This greatly reduces the number of parameters compared with giving every output pixel an independent connection to every input pixel.
Feature maps
Each learned filter produces a feature map indicating where its learned pattern is present. Stacking convolutional layers lets later features depend on progressively larger regions of the original input.
Stride and padding
Stride controls how far the kernel moves between positions. Padding controls how boundaries are treated and therefore affects output dimensions.
From local features to predictions
Convolutional layers are usually combined with nonlinear activations and later transformations that aggregate spatial information into task outputs.