Incident Management
A durable model for running production incidents: roles and span of control, severity that binds a commitment, mitigate-before-diagnose, communication cadence, and blameless review that actually changes the system rather than the person.
Incident management is the practice of detecting unplanned service degradation, coordinating a response, restoring service, and turning what happened into a change in the system. Its defining constraint is that it runs under time pressure with incomplete information, which is why almost all of it is about structure rather than cleverness: the structure is what lets a group of tired people make reasonable decisions quickly.
What an Incident Is
An incident is not "something broke". It is a condition affecting users that requires coordinated human attention right now. That definition does two things: it excludes the large class of alerts that are interesting but not urgent, and it includes degradations that no alert fired for — a customer report, a partner complaint, a number that looks wrong.
The single most valuable property of an incident process is that declaring is cheap. If declaring an incident means filling in a form, waking a director, and justifying yourself afterwards, engineers will wait, and the expensive part of every serious outage is the time between "something is wrong" and "we are working on it as a group". Make declaration a one-step action, make it acceptable to stand one down five minutes later, and accept a rate of false starts as the price of fast starts.
Roles and Span of Control
The failure this structure prevents is specific and common: eight engineers all debugging, nobody deciding, nobody talking to the business, and three people making conflicting changes to production at once.
| Role | Owns | Explicitly does not |
|---|---|---|
| Incident commander | Decisions, prioritisation, who does what, when to escalate | Debug, type commands, or own a workstream |
| Operations lead / SMEs | Investigating and applying changes in their area | Decide scope, or communicate externally |
| Communications lead | External and internal status updates on a cadence | Interrupt responders for detail mid-action |
| Scribe | Timeline of events, decisions and their times | Interpret or editorialise during the incident |
The hardest rule to hold is that the incident commander does not debug. The role is decision-making and coordination: who is doing what, what we are trying next, what we will do if it does not work, when we next update. The moment the commander drops into a terminal, the coordination stops and nobody notices for twenty minutes.
Span of control matters too. One commander can track a handful of parallel workstreams; beyond that, the answer is to delegate a sub-area to a lead rather than to track more threads. For long incidents, plan explicit handoffs — a commander at hour six is a liability, and a written handoff (current state, hypotheses tried and eliminated, actions in flight, open decisions) is what makes handing over safe.
Severity That Binds a Commitment
A severity scale is only useful if each level is defined by customer impact and binds a concrete commitment. Defining severity by component ("the database is down") fails because the same component failure can be invisible or catastrophic depending on what depends on it.
Each level should answer: who gets paged, how quickly must someone acknowledge, who is informed, how often do updates go out, and is a review mandatory. If a level does not change any of those answers, it is decoration. Two practical guards: publish examples of each level so the judgement is calibrated across teams, and allow severity to be raised or lowered mid-incident without ceremony — the first assessment is made with the least information.
The Lifecycle
Detect → an alert, a report, or a number that looks wrong
Declare → name it, open a channel, assign a commander
Triage → assess impact, set severity, page who is needed
Mitigate → restore service; capture evidence first if cheap
Monitor → confirm the symptom is gone, not just the alert
Resolve → stand down, hand off remaining cleanup as work
Review → blameless analysis; system changes with owners
The ordering that matters most is mitigate before diagnose. Restoring service and understanding the cause are different goals, and under time pressure they compete. A team that insists on understanding the failure before acting will keep users broken while it investigates. Stop the bleeding first; the evidence you need for the investigation can usually be captured before you act — take the heap dump, snapshot the logs, record the anomalous metric window — and then mitigate.
The highest-yield opening question is almost always what changed. Deploys, configuration pushes, feature flag flips, certificate rotations, schema migrations, dependency releases, traffic shifts and scheduled jobs are the usual candidates. A change timeline that a responder can read in one place, correlated against the start of the symptom, shortens more incidents than any amount of intuition.
Mitigation Levers, Ranked by Reversibility
Prefer the lever you can undo. Under uncertainty, a reversible action that might not help is better than an irreversible one that probably will.
- Roll back the most recent change. Highest yield, usually reversible, and it does not require understanding the failure.
- Turn off the feature flag. Faster than a deploy and narrower in blast radius, when the suspect path is flagged.
- Shed or throttle load. Serving a subset well beats serving everyone badly; reject cheaply at the edge rather than deep in the stack.
- Fail over or drain. Move traffic away from a failing zone, region or replica — only if you have evidence the target is healthy and has capacity.
- Scale out. Effective for genuine demand, useless or harmful when the bottleneck is a shared dependency or a lock.
- Restart. Often works, frequently destroys the evidence, and can trigger a thundering herd against cold caches and dependencies.
Two cautions. First, a mitigation can cause the next incident: a rollback across a schema change, a failover that lands full traffic on a cold replica, a scale-out that exhausts a connection pool downstream. Say the intended action out loud, and have someone check it against known dependencies before it is applied. Second, restarting to clear a symptom destroys the evidence — acceptable when service is down, costly when you are merely uncomfortable.
Communication
Separate the two audiences. The coordination channel is for responders and is dense, technical and noisy. The status channel is for everyone else and should be readable by someone with no context. Mixing them means either executives reading raw debugging or responders spending their attention on tone.
Update on a cadence, not on progress. An update that says "no change, still investigating, next update in thirty minutes" is valuable: silence is read as either "fixed" or "abandoned", and both readings generate interruptions that cost the responders more time than the update did. State impact in terms of what users cannot do, avoid speculation about cause, and never state a restoration time you cannot defend — a missed estimate costs more credibility than no estimate.
Review That Changes the System
A blameless review starts from the position that people acted reasonably given what they knew at the time, and that if a human error was sufficient to cause the outage, the system permitted it. This is not kindness; it is an accuracy requirement. A review that identifies a person as the cause stops at the point where the useful information starts, and it guarantees that the next person hides the next near-miss.
Some concrete markers of a review that is working:
- Counterfactuals ("should have noticed", "could have checked") are challenged rather than recorded — they describe a world that did not exist, not a change you can make.
- The narrative includes what responders believed at each point, not just what was true.
- Contributing factors are plural. Complex systems rarely fail from one cause; a single root cause in the document usually means the analysis stopped early.
- Things that went well are recorded, because those are the defences you must not accidentally remove later.
- Near-misses get reviews too, at lower cost — they carry the same information without the damage.
Action items need owners, a stated bar for "done", and a place they are actually tracked alongside other work. The most common quiet failure of incident programmes is an action-item graveyard: reviews are written, items are filed, nothing is scheduled, and the same incident recurs with a better-written document.
Failure Modes
| Failure mode | Why it happens | Counter |
|---|---|---|
| Heroism | One expert who always fixes it, so nothing is documented | Rotate the fixer; make runbooks a deliverable of the review |
| No handoff on long incidents | Nobody wants to hand over mid-crisis | Scheduled handoffs with a written state transfer |
| Severity inflation | Every incident is top severity, so nothing is | Impact-based definitions with published examples |
| Silent status page | Responders are busy; comms is nobody's job | A named communications role, updating on a clock |
| Action-item graveyard | Items filed but never scheduled | Owners, due dates, and the same tracker as feature work |
| Alert fatigue | Pages for causes and for conditions nobody acts on | Page only on symptoms needing a human now; the rest are tickets |
| Fix verified by absence of alert | The alert cleared before the symptom did | Verify against the user-visible signal, not the alert state |
Incident management is downstream of two other practices. It is only as fast as your ability to see the system, which is the job of observability; and it is only as rehearsed as your failure injection, which is the job of chaos engineering and game days. If the first time a runbook is executed is during a real outage, you are testing the runbook and the service at the same time.