Reliability

Stress-test an uptime budget before teams spend it

Stress-test an uptime budget by proving how downtime is measured, how burn becomes visible, and who can stop further risk when the remaining budget is thin.

Uptime budgets often quote a percentage while leaving the clock, the excluded maintenance windows, and the customer-facing consequence undefined. A 99.9 percent target can still allow painful outages if partial degradation is not counted. The exercise should follow the actual service definition and recent incident timestamps rather than a round number on a slide.

Map the measurement boundary

List the services covered, success criteria, polling or synthetic checks used, regional scope, and exclusions. Mark whether customer-reported failures count when monitors stay green. Attach ninety days of availability calculations, dependency maps, and the burn from each major incident. Include the remedy customers receive when the budget is exhausted.

State which changes consume budget intentionally, such as risky migrations, and which events are treated as external. For each intentional burn, define a measurable allocation and the action taken when remaining budget drops below a named threshold. Write the formula used to convert minutes of impact into budget burn so two teams can produce the same number.

Use scenarios with visible consequences

  1. A dependency outage is excluded in the contract while customers experience a full product stop.
  2. Partial errors stay under the failure threshold but destroy a peak-hour workflow.
  3. A planned maintenance window overruns and overlaps a customer commitment.
  4. Two regions disagree on status while the global budget continues to look healthy.
  5. An engineering team schedules a cutover that would consume the remaining quarterly budget in one evening.

For every scenario, ask reliability to show the detection signal and the burn accounting entry. Have product define acceptable customer messaging. Legal or customer success should identify any remedy that the budget implies but finance has not funded. Support should see the language customers will receive.

Judge whether the budget is operable

Run the policy and evidence through Pingpong. Reject answers that depend on informal remaining headroom without a measured remaining balance and an owner who can block further risk. Test whether dashboards distinguish monitor gaps from true availability, and whether a freeze decision can be made without assembling an ad hoc committee.

Related drills include an uptime commitment, an SLA change, and the SRE manager seat. The war-game decisions hub lists adjacent reviews.

Before publishing the budget, replay one historical incident through the accounting rules, capture the resulting burn, and fail the drill if any stakeholder needs an undocumented exception to explain the number.