Learning path

Full curriculum

Full curriculum

Unit content

Parallel granularity and overhead

Parallel work must be divided finely enough to expose concurrency but coarsely enough that useful computation dominates management cost.

The granularity of a task is the amount of useful work it performs before synchronizing, communicating or returning to the scheduler.

If one task performs only $1\ \mu s$ of useful work but creating and scheduling it costs $2\ \mu s$, parallelizing that unit makes the overhead larger than the computation itself.

At the opposite extreme, splitting a one-second job into only two half-second tasks cannot keep eight available cores busy.

Practical implementations therefore combine small logical iterations into chunks or stop recursive subdivision below a cutoff size. The best granularity depends on task cost, processor count, cache behavior, scheduling overhead and variation between tasks.

Granularity also affects responsiveness and load balance: smaller tasks are easier to redistribute, while larger tasks reduce scheduling traffic.

Parallel decomposition is therefore not simply “make as many tasks as possible”. It is a trade-off between exposing enough ready work and amortizing the cost of managing that work.