Unit content
Pipeline parallelism and throughput
Pipeline parallelism divides processing into stages and overlaps different items across those stages.
Suppose each item passes through three stages taking 2 ms, 5 ms and 3 ms. Processing one item still has at least
$$2+5+3=10\text{ ms}$$
of pipeline latency. But once the pipeline is full, different items can occupy the three stages simultaneously.
The steady-state throughput is limited by the slowest stage. Here the 5 ms stage can accept at most about
$$\frac{1}{0.005}=200$$
items per second, even if the other stages are faster.
Queues between stages absorb short timing differences, but an indefinitely slower consumer causes backpressure: work accumulates unless production slows, items are dropped, or capacity is increased.
Pipeline parallelism differs from data parallelism. Data parallelism replicates similar work over different data; a pipeline overlaps different kinds of work on different items.
Balanced stages improve throughput, while excessively deep pipelines can increase latency, buffering and coordination overhead.