Unit content
Incident response and post-incident learning
An incident is a period in which a system meaningfully fails to provide the behavior expected from it.
Incident response should first restore acceptable service. Typical priorities are:
- identify the user-visible impact;
- reduce or stop that impact through rollback, traffic changes, disabling a feature or another safe mitigation;
- preserve enough evidence to understand what happened;
- repair the underlying causes after immediate risk is controlled.
During an incident, a shared timeline and explicit ownership reduce duplicated or conflicting actions.
After recovery, a post-incident review reconstructs contributing conditions and identifies changes that would prevent recurrence or reduce impact. Useful actions change code, tests, deployment safeguards, observability, documentation or operational boundaries; assigning blame to one person does not explain why the system allowed one mistake to become an outage.
The goal is not to prove that no one erred. It is to turn a real failure into better defenses and faster future recovery.