War-game a prompt change review by confronting what the new instructions change in behavior, how regressions are measured, and who can reverse the change under load.
Prompt edits can look small while moving refusal boundaries, tool calls, tone, or factual caution. Diffs in a repository do not show which customer cohorts will notice first. The review needs the before and after prompts, the evaluation set that claims to cover the change, and a rollback path that does not depend on a redeploy nobody owns.
Build the review packet
Include the current prompt, the proposed prompt, intended behavior changes, prohibited behavior changes, tool schemas touched by the prompt, and the evaluation plan with sample counts by scenario. Attach recent production traces for the behaviors you claim to improve. Name the owner who can revert the prompt identifier in production.
Write one paragraph stating what should improve and what must not degrade. Keep latency, cost, safety, and product copy separate so each claim can be tested. State whether the change is global, cohort-gated, or experiment-backed. If the prompt references tools, include the exact tool schemas and any new permissions the change implies.
Run four pressure readings
Product owner: tests whether the intended user outcome actually moves on the evaluation set and in a small canary.
Eval owner: challenges coverage gaps, grader bias, and scenarios that only appear in live traffic.
Safety or trust reviewer: looks for refusal regressions, overconfident answers, and tool misuse introduced by new instructions.
Support lead: asks which new failure mode will become tickets and whether macros already exist for it.
Give every seat the same packet. Capture questions the packet cannot answer. Those gaps should become an explicit owner, a narrower rollout, or a hold.
Decide with a reversible path
Use Pingpong to simulate a canary where one cohort improves while another increases escalations, a tool-call rate spike, and a request to ship because a launch date is fixed. Ask reviewers which metric would stop expansion and where that metric is visible. Have ML ops confirm the rollback command against the actual serving config.
Pair this session with a model eval gate, the ML ops lead seat, or a feature flag rollout. The war-game decisions hub covers more product rooms.
Before promotion, require the prompt owner to answer the hardest safety failure using only the evaluation packet, then attach that answer to the release notes for the on-call rotation.