Stress-test a rollback window by proving that each recovery step can detect harm in time, reverse cleanly, and avoid trapping operators inside irreversible data or config moves.
Rollback plans often list a time target while leaving metric windows, migration direction, and abort authority implicit. Those edges decide whether a bad change stops inside the window or forces a forward fix under pressure. The exercise should follow actual services, dashboards, and on-call paths rather than a clean recovery slide.
Inventory the recovery path
List every step with its time budget, success metrics, abort thresholds, owners, and notification channels. Mark changes that cannot reverse without a data rewrite. Attach the last three rollbacks with raw timelines and any waivers. Include the source of truth for each metric during the observation window.
Define the phases for detect, decide, execute, verify, and communicate. Each phase needs an owner and an exit condition. Write the point after which rollback would require a different procedure, then review whether that action is still permitted. Capture maximum acceptable error and latency deltas in measurable units, including which customer cohorts are excluded from the gate and why.
Include the calendar of known release and migration events for the next two quarters: schema freezes, partner cutovers, and traffic peaks that shrink the usable window. A rollback budget that ignores those dates will look calm until the week they land.
Failure drills
- A minority cohort fails while the aggregate score stays inside the gate.
- A schema change has already written irreversible rows.
- The primary dashboard lags beyond the planned observation window.
- An operator expands forward because a partner deadline is close.
- A dependency degrades only under the rollback's traffic shape.
- Abort authority is unclear at 2 a.m. and the page lands on the wrong rotation.
For each drill, identify detection time, customer impact, containment, and the authority to pause. Require commands and dashboard links in the runbook. A statement that monitoring will catch it does not establish which alert fires or who receives it.
Prove recovery is timed and owned
Run the package in Pingpong with release manager, DevOps, SRE, and support seats. Ask product which user decision becomes unsafe first if the rollback is late. Ask SRE whether origin load can absorb a forced reverse. Ask DevOps to show the exact gate query used as the exit condition.
Related reviews include a hotfix criteria review, the release manager seat, and a canary percentage stress test. Browse the war-game decisions hub for adjacent controls.
Authorize the published window only after a timed drill restores the prior path inside the documented budget without an undocumented manual step.