Stress-test a DNS failover by proving that each TTL and health check can move traffic inside the planned window, survive resolver lag, and avoid trapping operators inside a green DNS slide that hides long-lived cache.
DNS failovers often list a minute count while leaving resolver behavior, health-check thresholds, and abort authority implicit. Those edges decide whether a late origin failure stops inside the window or leaves customers on a dead path. The exercise should follow actual records, probes, and on-call paths rather than a clean architecture deck.
Inventory the DNS path
List every record class with its TTL, health check, failover target, owners, and notification channels. Mark records that cannot reverse without a manual edit. Attach the last three failover incidents with raw timelines and any waivers. Include the source of truth for resolver-visible answers during the observation window.
Define the phases for schedule, cutover, observe, abort, and communicate. Each phase needs an owner and an exit condition. Write the point after which a stuck failover would require a different procedure, then review whether that action is still permitted. Capture maximum acceptable customer-visible lag in measurable units, including which cohorts are excluded from the new failover and why.
Include the calendar of known events for the next two quarters: DNS provider freezes, partner cutovers, and support peaks that shrink the usable change window. A failover budget that ignores those dates will look calm until the week they land.
Failure drills
- A minority money path keeps an old A record while the aggregate health dashboard stays green.
- A failover has already left a partner callback on a dead origin.
- The primary DNS dashboard lags beyond the planned observation window.
- An operator extends TTL because a launch demo is close.
- IPv4 and IPv6 answers collide under the new health threshold.
- Abort authority is unclear at 2 a.m. and the page lands on the wrong rotation.
For each drill, identify detection time, customer impact, containment, and the authority to force a DNS revert. Require commands and dashboard links in the runbook. A statement that monitoring will catch it does not establish which alert fires or who receives it.
Prove failovers are timed and owned
Run the package in Pingpong with network ops, SRE, support, and product seats. Ask product which user decision becomes unsafe first if DNS still points at a failed origin after the claimed window. Ask SRE whether abort can absorb a forced TTL extension. Ask network ops to show the exact record edit or provider API used as the exit condition.
Related reviews include the network ops lead seat, a CDN edge config review, and a CDN failover stress test. Browse the war-game decisions hub for adjacent controls.
Authorize the published failover only after a timed drill restores usable DNS answers inside the documented budget without an undocumented manual step.