Stress-test a read replica cutover before lag budgets, failover steps, and consistency claims harden into what every read path will quote.
Read replica cutovers fail when the doc invents lag precision the replicas never held, when traffic dual-counts the same query as primary and as replica, when apps still pin to retired hosts in private configs, and when platform cannot show who decides after a partial cut. A neat topology diagram is not evidence.
What belongs on the table
One sentence for why the cutover exists, which regions and services it covers, who owns lag monitors, failover, and break-glass, and the abort trigger if lag or error rates past a named threshold. Attach the replica inventory, sample lag metrics, app connection maps, and the measured path from announce to drained primary reads. If platform, SRE, and product disagree on which services are truly covered, stop and reconcile first.
Name the decision you will make if the stress test finds nothing new, and the delay criteria if any money-path service still lacks a named cutover owner or a verified rollback drill.
Failure modes worth seating
- Lag fiction: budgets that look tight while secondary metrics still skip weekends.
- Consistency blur: "eventual is fine" claims that invent completeness the last drill never showed.
- Config theater: connection strings that still host retired replica hosts.
- Break-glass lag: emergency paths that trail the customer-visible outage clock.
- Partial-cut silence: failures that land without a decision owner or measured lag.
Optional finance seat if write-path revenue reports bind the form. Optional support seat if customer-visible stale reads bind the form.
How to run the test
Feed Pingpong the draft cutover plan, drill notes, and open risk list. Early passes steelman the topology design. Later passes attack from platform, SRE, product, support, and skeptic seats. End with a pass that turns surviving objections into clearer owners, a timed rollback drill, or a hold. Delete invented "we already cut over cleanly" claims and dual-counted success rates.
Ask platform and product seats to price the behavior the published cutover will invite. If day-one docs promise zero stale reads while the last drill stranded billing dashboards for hours, buyers will treat the plan as false. Write the intended lag budgets, the failover checks, and the language you will refuse, then attack whether trust still holds under that discipline.
When the cutover coincides with a multi-region move or a schema migration, force platform and SRE seats to map every claim that still assumes last quarter's replica topology. Admin panels, reporting jobs, and partner embeds count. A read replica cutover that looks clean in a PDF while a critical path still pins to a retired host will fail on the first traffic wave. Related: stress-test a multi-region cutover, stress-test a schema migration, stress-test a CDN failover, pretend you are the CTO, and the war-game decisions hub. Process: how to run a Pingpong.