Infrastructure

Stress-test a CDN failover before you rely on it

Stress-test a CDN failover before origin maps, DNS cutovers, and runbook steps harden into what every edge outage will quote.

CDN failovers fail when the runbook invents cutover speed DNS never had, when cache keys dual-count the same object as purged and as live, when origins still lack health checks on money paths, and when ops cannot show who owns the decision after a partial regional failure. A neat architecture diagram is not evidence.

What to put on the table

One sentence for why the failover exists, which regions and object classes it covers, who owns DNS, origin health, and purge tooling, and the rollback trigger if error rates or TTL lag past a named threshold. Attach the origin map, sample purge logs, DNS design, and the measured path from detection to restored traffic. If platform, SRE, and product disagree on which paths are truly covered, stop and reconcile first.

Name the decision you will make if the stress test finds nothing new, and the delay criteria if any money path still lacks a named failover owner or a verified drill.

Attack surfaces

  • DNS fiction: TTLs that look short while resolvers still cache past the drill window.
  • Origin blur: health checks that still say "best effort" without a named threshold.
  • Cache theater: purge claims that invent completeness the edge never showed.
  • Runbook lag: steps that trail the customer-visible error rate.
  • Partial-region silence: failures that land without a decision owner or measured lag.

Optional security seat if certificate or WAF cutovers bind the form. Optional support seat if status updates bind the form.

How to run it

Feed Pingpong the draft failover plan, drill notes, and open risk list. Early passes steelman the design. Later passes attack from SRE, platform, product, support, and skeptic seats. End with a pass that turns surviving objections into clearer owners, a timed regional drill, or a hold. Delete invented "we already fail over cleanly" claims and dual-counted success rates.

Ask SRE and product seats to price the behavior the published failover will invite. If day-one runbooks promise five-minute recovery while the last drill stranded checkout for twenty, buyers will treat the plan as false. Write the intended regions, the health checks, and the language you will refuse, then attack whether trust still holds under that discipline.

When the failover coincides with a status page change or a new origin region, force eng and support seats to map every claim that still assumes the old edge layout. Partner embeds, mobile assets, and admin panels count. A failover that looks clean in a PDF while a critical asset still pins to one region will fail on the first outage wave. Related: stress-test a status page update, stress-test a migration plan, pretend you are the CTO, pretend you are the ops lead, and the war-game decisions hub. Process: how to run a Pingpong.