Unit content
SIMD vector processing
SIMD—single instruction, multiple data—uses one instruction to perform the same operation on several data elements packed into vector lanes.
If a processor has a vector instruction that adds four floating-point values at once, a loop such as
for i = 0 .. n-1:
y[i] = a[i] + b[i]
can process elements in groups of four instead of issuing one scalar addition per element.
SIMD works best when neighboring elements follow the same control flow and are laid out so the processor can load and store vectors efficiently. Compilers can sometimes auto-vectorize suitable loops; programmers can also use explicit vector operations or intrinsics.
Conditional behavior can be expressed with masks, enabling only selected lanes. Heavy per-element branching, irregular memory access or dependencies between neighboring iterations reduce the benefit.
The vector width is a hardware property, but portable code should express data-parallel intent rather than assume one fixed number of lanes.
SIMD is parallelism inside a processor core. It complements thread-level multicore parallelism: several cores can each execute vector instructions over different portions of the same dataset.