Product

War-game an experimentation policy before you freeze the rules

War-game an experimentation policy before soft guardrails, heroic sample sizes, and quiet carve-outs quietly become the company default.

Experimentation policies fail when ethics review is optional, when holdouts vanish under shipping pressure, when legal never sees user-facing variants, and when "temporary" flags live for quarters. A clean Notion page is not evidence that the org can absorb a bad experiment Monday.

Freeze one package

Write the scope of experiments you will allow, the review ladder you will defend, the decision rights you will honor, and the outcomes that still must hold if a test hurts trust. Then write the owner for each gate and the kill criteria if the policy finds missing consent or weak metrics. If you cannot name the owner or the kill criteria, the policy is decoration.

Freeze the draft rules, metric definitions, review calendar, and the customer language you would use if a test goes wrong. Attach residual risk you will accept. Split any secondary privacy rewrite into a separate decision. One policy under fire at a time.

Hostile seats

  • Product seat: shipping pressure that skips review and invents "obvious" wins.
  • Analytics seat: sample bias, metric drift, and soft causation dressed as proof.
  • Legal seat: variants that invent consent, dark patterns, or regulated claims.
  • Ethics seat: fairness gaps, vulnerable cohorts, and trust damage from surprise tests.
  • Skeptic seat: the exception that looks cleanest and has the least evidence.

Ask each seat for material objections, missing exhibits, and the smallest change that keeps the intent without the theater. Optional support seat if customer confusion sits inside the experiment blast radius.

Private pass

Run that package in Pingpong. Sequential review helps because an early pass can defend the learning story while later passes try to kill soft carve-outs and dual-counted spare capacity. Keep a change log of rules cut, sequenced, or clarified. If survival requires silence about a single hero experiment or a heroic uplift, stop. Rewrite the policy or delay the "approved" label.

Publish the activation calendar inside the package: when reviewers are trained, when flags age out, and when analytics locks definitions. An experimentation policy without dated owners is a rumor with a template. Ask product and analytics seats to break that calendar before anyone schedules the next "ship freely" note.

Related: war-game a growth experiment, war-game a pricing experiment, pretend you are the analyst, stress-test a feature flag rollout, and AI for high-stakes decisions. Mechanics: how it works.