Learning path

Full curriculum

Full curriculum

Unit content

Timeouts, retries and duplicate requests

When a distributed request does not produce a response before a deadline, the caller faces an ambiguous outcome. The remote operation may not have run, may still be running, or may have completed successfully while its response was lost.

A retry can recover from transient loss, but it can also execute the operation again.

For a read this may be harmless. For a command such as “charge this card” or “append this record,” blindly retrying can duplicate an effect.

Timeouts should therefore be chosen as control decisions rather than treated as failure detectors. Retries need bounds, backoff or other policies so an overloaded service is not flooded with synchronized repeated work.

The key distributed-systems question is not simply whether to retry, but what repeated execution means for the operation.