Unit content
NUMA memory locality
In a non-uniform memory access (NUMA) system, all processors may share one address space while memory access cost depends on which processor is physically closest to the data.
A multi-socket server commonly attaches part of main memory to each socket. A core can access remote memory through an interconnect, but that access usually has higher latency and consumes inter-socket bandwidth.
This makes thread placement and data placement interact. If one worker repeatedly processes a large array located near another socket, the program can become limited by remote-memory traffic even though the algorithm has no logical sharing.
Many systems use a first-touch policy: physical pages are placed near the processor that first writes them. Initializing each partition on the worker that will later process it can therefore improve locality.
Partitioning should aim to keep a worker's frequently used data local while limiting the amount of data shared across NUMA nodes. Migration can help changing workloads but also has cost.
NUMA turns locality from a cache-level concern into a machine-level placement problem: the same address is accessible everywhere, but not equally cheaply from everywhere.