Unit content
Distributed systems and partial failure
A distributed system is a set of independent processes or machines that cooperate by exchanging messages.
The defining difficulty is not merely that communication is slower than a local function call. Components can fail independently.
One node may be running while another has crashed. A message may be delayed, lost or duplicated. A network path may fail while both endpoints remain healthy.
This creates partial failure: one participant cannot always tell whether another participant is dead, disconnected or simply slow.
A timeout therefore provides evidence about elapsed time, not proof of remote failure.
Distributed algorithms must preserve their required properties despite these ambiguous outcomes. The design must state which failures it tolerates rather than assuming the whole system either works or fails together.