Stress-test a feature flag rollout before percent ramps, dark launches, and sticky defaults hide risk from the people who will own the incident.
Flag rollouts fail when kill switches are theoretical, when metrics lag the blast radius, when sticky cohorts cannot be pulled back cleanly, and when support learns about the change from customers. A ramp schedule in a ticket is not evidence of operational readiness.
What to put on the table
One sentence for what the flag changes, who sees it at each stage, the success and kill metrics, the rollback owner, and the communication plan for support and sales. Attach dependency maps, observability dashboards, prior incident notes, and the data or contract constraints that bind the ramp. If eng, product, and support disagree on blast radius, stop and reconcile first.
Name the decision you will make if the stress test finds nothing new, and the delay criteria if observability or rollback is not ready.
Attack surfaces
- Blast radius: which cohort or tenant cannot be isolated if the flag misbehaves.
- Observability: which failure mode has no alert until customers complain.
- Rollback honesty: sticky state, migrations, or caches that survive a flag flip.
- Support load: scripts and capacity versus claimed exposure.
- Compliance and contracts: features you promised would stay off for a segment.
Optional security seat if the flag touches auth, data export, or admin power. Optional finance seat if the flag changes billing behavior.
How to run it
Feed Pingpong the rollout memo and exhibits. Early passes steelman the ramp. Later passes attack from eng ops, support, customer, and skeptic seats. End with a pass that turns surviving objections into smaller cohorts, clearer kill criteria, or a hold. Delete invented SLO numbers and dual-counted "successful" ramps that never tested rollback.
Force a day-after narrative: what on-call sees, which enterprise accounts escalate, and what happens if the flag is stuck on for a sticky cohort. If those stories are stronger than your rollback plan, fix the package before you raise exposure. Separate canary from broad ramp in the memo. If blended success criteria hide a soft stage, ask ops and product seats to attack until each stage has owners and kill criteria.
When the flag touches billing, entitlements, or data retention, force a counsel or finance seat into the loop before any broad ramp. Write the customer-facing explanation for a rollback into the package so support is not inventing language during an incident. Ask ops and support seats to attack that explanation until it matches what the flag can actually undo.
Related: war-game a product launch, stress-test a security incident response, before you launch, before you kill the feature, and the war-game decisions hub. Process: how it works.