Unit content
Parallel granularity and overhead
Parallel work must be divided finely enough to expose concurrency but coarsely enough that useful computation dominates management cost.
The granularity of a task is the amount of useful work it performs before synchronizing, communicating or returning to the scheduler.
If one task performs only $1\ \mu s$ of useful work but creating and scheduling it costs $2\ \mu s$, parallelizing that unit makes the overhead larger than the computation itself.
At the opposite extreme, splitting a one-second job into only two half-second tasks cannot keep eight available cores busy.
Practical implementations therefore combine small logical iterations into chunks or stop recursive subdivision below a cutoff size. The best granularity depends on task cost, processor count, cache behavior, scheduling overhead and variation between tasks.
Granularity also affects responsiveness and load balance: smaller tasks are easier to redistribute, while larger tasks reduce scheduling traffic.
Parallel decomposition is therefore not simply “make as many tasks as possible”. It is a trade-off between exposing enough ready work and amortizing the cost of managing that work.