Unit content
Data parallelism and parallel loops
Data parallelism applies the same operation to many elements or regions of a dataset at the same time.
A loop is parallelizable when iterations do not have conflicting dependencies. For
for i = 0 .. n-1:
y[i] = 2 * x[i] + 1
each iteration reads and writes different elements, so the iteration space can be divided among workers without changing the result.
A parallel loop assigns subsets of iterations to different execution resources. The partition may be contiguous blocks, cyclic chunks, tiles, or dynamically assigned work depending on the access pattern and cost of each iteration.
Not every element-wise-looking loop is independent. In
for i = 1 .. n-1:
x[i] = x[i-1] + 1
iteration $i$ depends on the result of iteration $i-1$, so naive simultaneous execution changes the computation.
Data parallelism is especially effective when many iterations perform similar amounts of work over regularly laid-out data. It underlies vector instructions, GPU kernels, image processing, dense numerical operations and many machine-learning workloads.