Learning path

Full curriculum

Full curriculum

Arrows go from each prerequisite to the units that depend on it. Hover or focus a unit to highlight its path.

Unit content

Pipeline parallelism and throughput

Pipeline parallelism divides processing into stages and overlaps different items across those stages.

Suppose each item passes through three stages taking 2 ms, 5 ms and 3 ms. Processing one item still has at least

$$2+5+3=10\text{ ms}$$

of pipeline latency. But once the pipeline is full, different items can occupy the three stages simultaneously.

The steady-state throughput is limited by the slowest stage. Here the 5 ms stage can accept at most about

$$\frac{1}{0.005}=200$$

items per second, even if the other stages are faster.

Queues between stages absorb short timing differences, but an indefinitely slower consumer causes backpressure: work accumulates unless production slows, items are dropped, or capacity is increased.

Pipeline parallelism differs from data parallelism. Data parallelism replicates similar work over different data; a pipeline overlaps different kinds of work on different items.

Balanced stages improve throughput, while excessively deep pipelines can increase latency, buffering and coordination overhead.