Ops

Stress-test an on-call rotation change before you publish it

Stress-test an on-call rotation change before the next incident proves the new schedule leaves a gap, burns out a thin bench, or breaks an SLA customers already paid for.

Rotation changes fail when handoff rules are fuzzy, when secondary coverage is theater, when PTO and timezone reality are ignored, and when the paging policy still assumes the old team shape. A calendar screenshot is not evidence that the system can absorb a night of cascading alerts.

What to put on the table

One sentence for what changes, who is in scope, the effective date, the success metric after thirty days, and the rollback trigger. Attach the old and new rotation, escalation ladder, SLO or SLA commitments, and the last quarter of incident load by severity. If eng, support, and customer success disagree on who owns first response, stop and reconcile first.

Name the decision you will make if the stress test finds nothing new, and the delay criteria if coverage or training is not ready.

Attack surfaces

  • Coverage truth: nights, weekends, and timezone holes the new chart still leaves open.
  • Burnout math: pages per person under realistic spike load, not the average week.
  • Skill match: runbooks and services that only one person can still own.
  • Customer promises: SLA clocks that the new ladder cannot meet.
  • Handoff and paging: who gets woken, who acknowledges, and what "resolved" means.

Optional finance seat if contractor or overtime spend changes. Optional people seat if the change interacts with a freeze or attrition risk.

How to run it

Feed Pingpong the rotation draft, escalation ladder, and incident history. Early passes steelman the change. Later passes attack from primary on-call, secondary, customer, and skeptic seats. End with a pass that turns surviving objections into a phased rollout, clearer runbooks, or a hold. Delete invented "someone will catch it" coverage and dual-counted headcount.

When the change adds a new service or region, force eng and support seats to map every paging path that still assumes the old owner. Runbooks, chat aliases, and vendor escalations count. A schedule that looks clean in a spreadsheet while the pager still rings the wrong person will fail on the first real night.

Force a day-after narrative: what happens if two sev-1s overlap, if a key person is on PTO, or if a customer quotes the SLA during an outage. If those stories are stronger than your mitigation plan, fix the package before publish. Related: stress-test an SLA change, war-game a support escalation ladder, stress-test a security incident response, pretend you are the CTO, and the war-game decisions hub. Process: how to run a Pingpong.