Learning path

Full curriculum

Full curriculum

Arrows go from each prerequisite to the units that depend on it. Hover or focus a unit to highlight its path.

Unit content

Arithmetic intensity and memory-bandwidth limits

A parallel processor can execute arithmetic only as fast as data reaches its execution units.

The arithmetic intensity of a workload is the amount of computation performed per unit of data transferred from a limiting memory level, commonly measured in operations per byte:

$$I=\frac{\text{operations}}{\text{bytes transferred}}.$$

If memory bandwidth is $B$ bytes/s, data movement alone limits performance to approximately

$$P\le BI.$$

The processor also has a finite peak compute rate $P_{max}$, so a simple roofline bound is

$$P\le\min(P_{max},BI).$$

A low-intensity operation such as adding two large arrays performs little arithmetic for every byte loaded and stored, so adding more arithmetic units may not help once memory bandwidth is saturated.

A higher-intensity algorithm reuses data from registers or caches for many operations before fetching new data. Blocking matrix operations is a classic example.

This distinction explains why parallel speedup can flatten even when there are idle compute units: the workload may be memory-bound rather than compute-bound. Improving locality or reducing data movement can then matter more than adding threads.