Unit content
Heterogeneous CPU-GPU memory and data transfer
A CPU and an accelerator such as a GPU may execute different parts of one program while having different access costs—or even different address spaces—for memory.
In an explicit host/device model, data needed by a GPU kernel must be copied from host memory to device memory before execution, and results copied back afterward.
For a computation taking $5$ ms on the GPU but requiring $8$ ms to copy inputs and $8$ ms to copy results, acceleration loses to a $15$ ms CPU implementation:
$$T_{GPU}=8+5+8=21\text{ ms}.$$
The useful arithmetic must therefore be large enough, repeated enough, or overlapped enough to amortize transfer costs.
Systems with shared or unified memory can make the same allocation addressable from both processors, but physical data may still migrate between memory domains. Easier addressing does not remove bandwidth, latency or locality costs.
High-performance heterogeneous programs keep data near the processor that will reuse it, batch work to reduce transfers, and overlap independent transfers with computation when the platform permits it.
Choosing where computation runs is therefore inseparable from choosing where its data resides.