Infrastructure

Stress-test a blue-green cutover before you rely on it

Stress-test a blue-green cutover before drain budgets, rollback steps, and parity claims harden into what every release path will quote.

Blue-green cutovers fail when the doc invents drain precision the last release never held, when traffic dual-counts the same request as blue and as green, when apps still pin to retired hosts in private configs, and when platform cannot show who decides after a partial flip. A neat topology diagram is not evidence.

What belongs on the table

One sentence for why the cutover exists, which services and regions it covers, who owns health checks, traffic flip, and break-glass, and the abort trigger if error rates or lag past a named threshold. Attach the environment inventory, sample flip metrics, app connection maps, and the measured path from announce to drained blue. If platform, SRE, and product disagree on which services are truly covered, stop and reconcile first.

Name the decision you will make if the stress test finds nothing new, and the delay criteria if any money-path service still lacks a named flip owner or a verified rollback drill.

Failure modes worth seating

  • Drain fiction: budgets that look tight while sticky sessions still skip weekends.
  • Parity blur: "green matches blue" claims that invent completeness the last drill never showed.
  • Config theater: connection strings that still host retired blue hosts.
  • Break-glass lag: emergency paths that trail the customer-visible outage clock.
  • Partial-flip silence: failures that land without a decision owner or measured error rate.

Optional finance seat if write-path revenue reports bind the form. Optional support seat if customer-visible stale reads bind the form.

How to run the test

Feed Pingpong the draft cutover plan, drill notes, and open risk list. Early passes steelman the flip design. Later passes attack from platform, SRE, product, support, and skeptic seats. End with a pass that turns surviving objections into clearer owners, a timed rollback drill, or a hold. Delete invented "we already flip cleanly" claims and dual-counted success rates.

Ask platform and product seats to price the behavior the published cutover will invite. If day-one docs promise zero stale sessions while the last drill stranded billing dashboards for hours, buyers will treat the plan as false. Write the intended drain budgets, the flip checks, and the language you will refuse, then attack whether trust still holds under that discipline.

When the cutover coincides with a schema migration or a feature flag rollout, force platform and SRE seats to map every claim that still assumes last quarter's blue-green topology. Admin panels, reporting jobs, and partner embeds count. A blue-green cutover that looks clean in a PDF while a critical path still pins to a retired host will fail on the first traffic wave. Related: stress-test a read replica cutover, stress-test a feature flag rollout, stress-test a multi-region cutover, pretend you are the SRE manager, and the war-game decisions hub. Process: how to run a Pingpong.