Learning path

Full curriculum

Full curriculum

Unit content

CPU cache hierarchy and cache lines

Processors execute instructions much faster than main memory can supply arbitrary data. CPU caches keep recently used memory blocks close to the processor so that common accesses can be served with much lower latency.

Cache hierarchy

Modern processors commonly use several cache levels. Small L1 caches are closest and fastest; larger L2 and L3 caches provide more capacity at higher latency before an access must reach main memory.

Cache lines

Memory is transferred between cache levels in fixed-size blocks called cache lines. Reading one byte can therefore bring neighboring bytes into the cache as well.

This makes the physical layout of data important: accessing one useful value may also fetch nearby values that will either be used soon or occupy cache space unnecessarily.

Hits and misses

A cache hit occurs when the needed line is already present at the required cache level. A cache miss requires fetching it from a slower level.

Working sets

If the actively used data fits in a fast cache, repeated accesses can be inexpensive. When the working set exceeds cache capacity or accesses continually replace one another, performance can fall sharply.

Caches are mostly transparent to program semantics, but they strongly affect the real cost of different memory-access patterns.