Storage operations

Stress-test a restore drill SLA before cutover fiction outlives the drill

Stress-test a restore drill SLA by proving that each class and vault can return usable data inside the planned window, survive tool lag, and avoid trapping operators inside a green restore slide that hides long-lived failure.

Restore SLAs often list a minute count while leaving vault behavior, integrity checks, and abort authority implicit. Those edges decide whether a late backup failure stops inside the window or leaves customers on empty data. The exercise should follow actual restore jobs, probes, and on-call paths rather than a clean architecture deck.

Inventory the restore path

List every class with its RTO, vault, restore target, owners, and notification channels. Mark classes that cannot reverse without a manual rebuild. Attach the last three restore drills with raw timelines and any waivers. Include the source of truth for customer-visible data during the observation window.

Define the phases for schedule, restore, observe, abort, and communicate. Each phase needs an owner and an exit condition. Write the point after which a stuck restore would require a different procedure, then review whether that action is still permitted. Capture maximum acceptable customer-visible lag in measurable units, including which cohorts are excluded from the new SLA and why.

Include the calendar of known events for the next two quarters: vault freezes, partner cutovers, and support peaks that shrink the usable drill window. An SLA budget that ignores those dates will look calm until the week they land.

Failure drills

  1. A minority money path keeps an old backup while the aggregate restore dashboard stays green.
  2. A drill has already left a partner callback on empty data.
  3. The primary vault dashboard lags beyond the planned observation window.
  4. An operator extends RTO because a launch demo is close.
  5. Hot and cold restores collide under the new integrity threshold.
  6. Abort authority is unclear at 2 a.m. and the page lands on the wrong rotation.

For each drill, identify detection time, customer impact, containment, and the authority to force a restore abort. Require commands and dashboard links in the runbook. A statement that monitoring will catch it does not establish which alert fires or who receives it.

Prove restores are timed and owned

Run the package in Pingpong with storage ops, SRE, support, and product seats. Ask product which user decision becomes unsafe first if restores still fail after the claimed window. Ask SRE whether abort can absorb a forced RTO extension. Ask storage ops to show the exact restore job or provider API used as the exit condition.

Related reviews include the storage ops lead seat, a backup retention ladder review, and a backup restore drill stress test. Browse the war-game decisions hub for adjacent controls.

Authorize the published SLA only after a timed drill restores usable data inside the documented budget without an undocumented manual step.