Unit content
Grouped and depthwise-separable convolutions
A standard convolution mixes spatial information and channel information in one operation. Grouped and depthwise-separable convolutions reduce that coupling to lower computational cost.
Grouped convolution
Input and output channels are partitioned into groups. Each output channel only uses input channels from its own group.
With one group, the operation is an ordinary convolution. Increasing the number of groups reduces connections and computation.
Depthwise convolution
A depthwise convolution applies a separate spatial filter to each input channel instead of mixing channels together.
Spatial pattern extraction and channel mixing are therefore separated.
Pointwise convolution
A $1\times1$ pointwise convolution mixes channel values at each spatial location without looking across neighboring positions.
Combining depthwise spatial filtering with pointwise channel mixing gives a depthwise-separable convolution.
Trade-off
These factorizations can greatly reduce parameters and multiply-add operations, especially in mobile or real-time models. The price is a more constrained interaction pattern, so efficiency does not automatically imply equal accuracy for every task.