War-game a runbook ownership map by checking whether every critical path has one accountable owner, current links, and a reachable backup when the primary is offline.
Ownership maps look complete when every service has a cell filled. They fail when two teams share a step without a decision rule, when links point at retired dashboards, and when the backup pager is a mailing list nobody monitors on weekends. The review should start from the last three incidents and the actual steps people followed, then compare that trail to the published map.
Build the map from incidents, not org charts
For each money-path or high-urgency service, list detection, triage, mitigation, customer communication, and post-incident follow-up. Name a primary owner, a backup, the escalation path, and the systems each person must access. Attach on-call schedules, access grants, and the date each runbook link was last verified. If a step requires a vendor console, document who holds the credential and how a substitute gets it.
Mark steps that currently depend on tribal knowledge. Those become either written procedures with owners or explicit risks with residual acceptance. Do not leave them as "someone will know."
Attack the map with concrete cases
- Primary owner is unreachable; backup lacks console access for the mitigation step.
- Two teams both believe they own customer status updates and post conflicting messages.
- A critical dashboard link returns 404 during the first fifteen minutes.
- A dependency outage lands outside any named service on the map.
- Weekend coverage routes to a rotation that no longer includes the owning team.
- A "temporary" manual step has lived in the runbook for two quarters without an owner for automation.
For each case, record detection lag, who decides, and what changes in the map. If the answer is a hallway conversation, rewrite the ownership cell before the next freeze or canary window.
Close with named gaps only
Run the session in Pingpong with SRE, DevOps, support, and a product owner seat. Ask support which customer-facing promise breaks first when ownership is unclear. Ask DevOps whether deploy freezes and exception paths still point at living people. Ask SRE to confirm that every severity level has a reachable decision owner.
Pair with the SRE manager seat, an incident severity matrix stress test, and a deploy freeze exception review. More formats sit in the war-game decisions hub.
Adopt the map only after every high-urgency row lists one primary, one backup, a verified link set, and a dated access check for the tools those people need.