Unit content
Arithmetic intensity and memory-bandwidth limits
A parallel processor can execute arithmetic only as fast as data reaches its execution units.
The arithmetic intensity of a workload is the amount of computation performed per unit of data transferred from a limiting memory level, commonly measured in operations per byte:
$$I=\frac{\text{operations}}{\text{bytes transferred}}.$$
If memory bandwidth is $B$ bytes/s, data movement alone limits performance to approximately
$$P\le BI.$$
The processor also has a finite peak compute rate $P_{max}$, so a simple roofline bound is
$$P\le\min(P_{max},BI).$$
A low-intensity operation such as adding two large arrays performs little arithmetic for every byte loaded and stored, so adding more arithmetic units may not help once memory bandwidth is saturated.
A higher-intensity algorithm reuses data from registers or caches for many operations before fetching new data. Blocking matrix operations is a classic example.
This distinction explains why parallel speedup can flatten even when there are idle compute units: the workload may be memory-bound rather than compute-bound. Improving locality or reducing data movement can then matter more than adding threads.