Release operations

War-game a deploy freeze exception before it becomes routine

War-game a deploy freeze exception by testing whether the exception criteria, blast radius, and rollback proof still protect customers when a deadline is close and the change looks small.

Freeze policies fail when "customer critical" is undefined, when exceptions dual-count the same change as urgent and as low risk, and when rollback drills live only as a promise. The review needs one written definition of an allowed exception, one owner who can deny it, and a documented consequence when the exception fails.

Freeze the exception proposal

Write the freeze window, services covered, exception categories, required evidence, approvers, notification list, and rollback drill requirements. Attach the last five exceptions with their outcomes and any customer-visible impact. Identify exclusions in plain language. If a class of change is omitted from the freeze, show how that omission affects risk rather than leaving it as a footnote.

Name the approval choice and the conditions that force a hold. Include the ticket or form fields that compute eligibility so two reviewers can reproduce the same allow or deny label from the same packet.

Seat the freeze from both sides

Feature owner
Defends why the change cannot wait and what customer harm the delay would cause.
DevOps lead
Challenges artifact identity, promotion path, and whether the exception reuses an untested train.
SRE manager
Tests whether monitors and on-call coverage match the claimed blast radius.
Support lead
Shows which customer-facing symptoms appear first if the exception is wrong.
Skeptic
Finds the strongest urgency claim with the weakest rollback evidence.

Run pressure cases on the exception form

Use Pingpong to walk through a near-miss "tiny config" change, a partner deadline, a security patch that arrives mid-freeze, and a request to waive the rollback drill because staging already looked green. For each case, start from the documented freeze language. Ask who can expand the window and which evidence is required to reverse a deny. Any step that depends on an unnamed person becomes a release condition.

Ask the room to replay one historical exception under the proposed rules. If the historical case would have shipped while later causing a known incident, revise the policy before treating it as standard.

Compare the proposal with the DevOps lead seat and a canary percentage stress test. If ownership of recovery steps is unclear, review a runbook ownership map. More operating decisions live in the war-game decisions hub.

Publish the exception path only after a dry run can reproduce the same allow or deny label from the stored ticket fields without manual reinterpretation.