Role-play

Pretend you are the SRE manager

Pretend you are the SRE manager so lag budgets, on-call theater, and rollback blur fail before a reliability package absorbs them.

Optimistic reliability packages optimize for "ship now, monitor later." The SRE manager seat does the opposite. It asks which SLO invents error budget headroom the last quarter never held, which alert dual-counts the same page as covered and as quiet, which "required" rollback drill is already optional in practice, and which service still lacks a named owner. A neat dashboard screenshot is not evidence that next week's cutover will clear.

Cast with a real mandate

Name a real job: underwrite a cutover without inventing on-call capacity, clear a latency claim that platform can reconcile under load, or defend a freeze without dual-counted green deploys. Give constraints: the evidence standard for readiness, the services you will refuse to leave unowned, and the reliability claims you will not teach when break-glass paths are not ready. Without constraints the seat becomes cartoonish. With constraints it produces questions you might actually hear in an incident review, a capacity fight, or a dispute over who owns weekend pages.

Prompt example: "SRE manager: list the top reasons to delay this reliability package, the service with the weakest owner, the SLO claim that worries you most, and the ten diligence questions you would send after review. Stay inside a realistic mandate."

Keep these outputs

  • Top reasons to challenge, delay, or rewrite this reliability package.
  • The SLO or uptime claim that looks strongest and is least evidenced.
  • The service, region, or dependency that would break first under a forced ship.
  • What would make you accept residual reliability risk in writing.
  • The ten hardest follow-up questions after the meeting.

Run that brief in Pingpong against the real service map, on-call roster, drill notes, and open risk list. Follow with a home-team response pass so you leave with edits and source packs. When the decision is a public status claim or a customer-facing cutover, run this seat after platform and product attacks so it can use earlier objections as ammunition.

When the plan leans on a single pager rotation, a single "we will rollback in minutes" promise, or a single hero on-call, force the seat to price concentration risk in writing. Ask what happens if the runbook board still hosts retired owners, if SRE still lacks a named owner for multi-region failovers, or if product keeps teaching uptime language SRE already retired. Pair with pretend you are the CTO, stress-test an on-call rotation change, stress-test an incident postmortem, stress-test a read replica cutover, and the war-game decisions hub. See how to run a Pingpong.

If the package coincides with a multi-region move or a schema migration, ask the SRE manager seat to map every claim that still assumes last quarter's topology. Admin panels, partner embeds, and billing jobs count. A reliability memo that looks clean in a slide while a critical path still pins to a retired host will fail on the first traffic wave.