Role-play

Pretend you are the reliability ops lead

Pretend you are the reliability ops lead and force the reliability package to survive questions about who can rewrite an error budget, how a freeze is timed, and which on-call paths still lack a named owner when the live sample disagrees with the slide.

This seat sits between observability tools and the uptime claims leadership will quote. It asks which budget trees invent completeness the page export never measured, which freeze cuts look crisp on a slide and soft in the live sample, and what happens when a reliability exception lands because ownership is incomplete. A green SRE board does not answer those questions. The review needs the error-budget map, freeze calendar, exception packet, and the named person who can freeze a release path.

Hand the seat a usable brief

Hand over the observability tools in use, error-budget draft, freeze cut plan, exception ladder packet, last customer-visible incidents, and one recent case where a delayed freeze hurt a reliability narrative. Include which product promises still bypass the same review. State the decision: clear the reliability change, revise specific controls, or hold until a freeze owner is named.

Boundaries matter. The reliability ops lead can challenge untested budget trees, freeze windows that ignore measured recovery time, exceptions without a kill switch, and override paths without logging. Customer outcome ownership stays with support ops. Privilege ceilings belong in the same packet so a quiet dashboard does not hide a brittle on-call path.

Questions that must get numbers

  • Which critical service class still lacks a tested freeze path with a measured age-out, and who owns the gap?
  • What must hold before an error-budget cut can promote into a lasting rule, and who can waive it without a written reason?
  • How does a declined reliability exception become visible to the requester within the claimed window?
  • What is the measured time from a failed SLO finding to a human with freeze authority?
  • Which shared override can ship many release changes without failing a single reliability health check?
  • Who has authority to pause exceptions or force a reliability rollback at week end without waiting for the system owner?

Require observed, inferred, or unknown labels. Observed claims need a source. Unknowns become owners and due dates. When two teams claim the same reliability authority, force one named decision before the next tool change starts.

Convert objections into freeze gates

Run the role in Pingpong with the same exhibits the reliability team will use. Have the home team answer each objection in writing. The useful output is a short reliability ledger: approved service classes, blocked classes, freeze windows, and the person who can call a freeze.

For adjacent reliability work, pair this seat with the SRE manager seat, a release quality gate review, an on-call rotation stress test, and the quality ops lead seat. Incident narrative often needs an outage comms script stress test. The war-game decisions hub has more seats. Before approving the package, make the reliability ops lead write the exact freeze and budget check that will decide whether the change continues.