Infrastructure

Stress-test an incident severity matrix before you rely on it

Stress-test an incident severity matrix before severity labels, page rules, and ownership claims harden into what every incident path will quote.

Incident severity matrices fail when the doc invents triage precision the last incident never held, when pages dual-count the same alert as Sev-1 and as Sev-2, when on-call still pins to retired severity language in private runbooks, and when SRE cannot show who decides after a contested label. A neat severity table is not evidence.

Put the real matrix on the table

One sentence for why the matrix exists, which services and customer classes it covers, who owns severity calls, customer comms, and break-glass, and the abort trigger if page lag or reopen rates past a named threshold. Attach the current severity inventory, sample incident timelines, page metrics, and the measured path from detect to first customer update. If SRE, support, and product disagree on which incidents are truly covered, stop and reconcile first.

Name the decision you will make if the stress test finds nothing new, and the delay criteria if any money-path incident class still lacks a named owner or a verified drill.

Where matrices quietly fail

  • Severity fiction: labels that look crisp while weekend pages still skip the published path.
  • Ownership blur: "follow the matrix" claims that invent completeness the last drill never showed.
  • Runbook theater: severity tables that still host retired page rules.
  • Comms lag: customer updates that trail the customer-visible outage clock.
  • Contested-label silence: fights that land without a decision owner or measured reopen rate.

Optional legal seat if contractual severity SLAs bind the form. Optional support seat if customer-facing severity language binds the form.

How to pressure it

Feed Pingpong the draft matrix, drill notes, and open risk list. Early passes steelman the severity ladder. Later passes attack from SRE, support, product, legal, and skeptic seats. End with a pass that turns surviving objections into clearer owners, a timed drill, or a hold. Delete invented "we already triage cleanly" claims and dual-counted page rates.

Ask SRE and support seats to price the behavior the published matrix will invite. If day-one docs promise Sev-1 pages in minutes while the last drill stranded billing customers for hours, buyers will treat the plan as false. Write the intended severity labels, the page checks, and the language you will refuse, then attack whether trust still holds under that discipline.

When the matrix coincides with an on-call rotation change or a status page rewrite, force SRE and support seats to map every claim that still assumes last quarter's severity language. Partner embeds, enterprise SLAs, and regional support queues count. An incident severity matrix that looks clean in a PDF while a critical path still pins to a retired page rule will fail on the first outage wave. Related: stress-test an incident postmortem, stress-test an on-call rotation change, stress-test a security incident response, pretend you are the SRE manager, and the war-game decisions hub. Process: how to run a Pingpong.