Platform operations

War-game a read replica strategy before you publish it

War-game a read replica strategy before lag budgets, stale-read allowances, and failover ownership harden into how every service reads data.

Read replica strategies fail when lag targets live only as slides, when stale reads are treated as harmless for money paths, when failover scripts name a person who is not on-call, and when success is measured as replica count rather than a measured freshness under peak load. A tidy topology diagram is not evidence.

Freeze first

State the workloads that will read from replicas, the maximum lag each workload can tolerate, who owns lag alerts, the failover procedure and abort criteria, and the write path that must never divert to a replica by accident. Attach recent lag histograms, cutover drill notes, connection pool settings, and the kill criteria if freshness claims hide customer-visible inconsistency. If platform, application, and support disagree on which reads may be stale, reconcile before seating.

Name the decision you will make if the war game finds nothing new, and the hold criteria if any money path lacks a lag budget or a named abort owner.

Attack surfaces

  • Lag fiction: targets that look fine in quiet hours and break during campaigns.
  • Stale money: checkout, billing, or entitlement reads that treat lag as cosmetic.
  • Failover theater: runbooks that assume a human who is not on the rotation.
  • Pool blur: one connection pool for replica and primary that hides routing mistakes.
  • False capacity: replica count celebrated while lag and reconnect storms stay unmeasured.

Optional finance seat if billing reads depend on replica freshness. Optional support seat if stale reads already generate tickets.

How to run it

Feed Pingpong the draft strategy, lag evidence, and open risk list. Early passes steelman the routing plan. Later passes attack from application, database admin, SRE, and support seats. End with a pass that turns surviving objections into tighter lag budgets, clearer abort owners, or a hold. Delete invented "replicas are always seconds behind" claims and dual-counted capacity.

Ask database admin and SRE seats to price the customer experience the strategy will create. If marketing language promises real-time reads while lag budgets allow minutes of delay, customers will learn the promise is theater. Write the intended routing, the escalation path for lag breaches, and the reads you will refuse to send to a replica, then attack whether the plan still holds under that discipline.

Force a day-after narrative: what happens if a replica falls minutes behind during a sale, if the primary fails mid-cutover, or if a pool sends writes to a replica for ten minutes. If those stories are stronger than your mitigation plan, fix the strategy before the next traffic peak. Related: pretend you are the database admin, stress-test a read replica cutover, stress-test a vacuum window, and the war-game decisions hub. Process: how to run a Pingpong.