Unit content
GPU SIMT execution and branch divergence
GPUs obtain high throughput by running very large numbers of lightweight threads over data-parallel workloads.
A common execution model is SIMT—single instruction, multiple threads. Threads are grouped into hardware batches often called warps or wavefronts. Threads in one group normally execute the same instruction together on different data.
A kernel such as
i = global_thread_index
if i < n:
y[i] = a[i] * x[i] + y[i]
can launch enough threads to cover a large array. The GPU schedules many groups so that when one group waits for memory, others can execute.
If threads in the same group take different branches, the group experiences branch divergence. The hardware may execute each path separately while disabling lanes that do not belong to that path, reducing useful parallel work.
GPUs therefore favor workloads with abundant independent operations, similar control flow and regular memory access. They trade strong single-thread latency for aggregate throughput.
SIMT resembles SIMD in executing common operations across lanes, but its programming abstraction exposes many logical threads rather than one explicit vector instruction.