ML operations

War-game a model eval gate before it becomes the release rule

War-game a model eval gate by testing whether its metrics, slices, and waiver path still protect users when a release is late and the numbers look close.

A gate can look rigorous while measuring the wrong behavior. Aggregate accuracy may hide a failing cohort. A holdout set may share an upstream feature pipeline with training. A human evaluation rubric may reward fluency while missing factual misses that matter in production. The review needs one definition of pass, one population that must be covered, and a documented consequence when the gate fails.

Freeze the gate proposal

Write the offline metrics, online proxies, slice definitions, sample sizes, statistical thresholds, data windows, and ownership for each check. Attach the last three promotions with their raw scorecards, any waivers granted, and the outcome after release. Identify exclusions in plain language. If a class of traffic is omitted from evaluation, show how that omission affects the claim rather than leaving it as a footnote.

Name the approval choice and the conditions that force a hold. This keeps the session from drifting into a general debate about model quality. Include the tool or script that computes the score so two reviewers can reproduce the same label from the same artifacts.

Seat the gate from both sides

Model owner
Defends what the candidate can demonstrate on the stated evaluation package.
Eval owner
Challenges leakage, slice coverage, metric substitution, and score stability across runs.
Product owner
Shows which user decisions become unsafe when a silent failure mode ships.
ML ops lead
Tests whether online monitors can detect the same failure the gate claims to catch.
Skeptic
Finds the strongest published claim with the weakest measurement path.

Run pressure cases on the scorecard

Use Pingpong to walk through a near-miss aggregate score, a failing minority slice, a contaminated holdout, and a request to waive because a partner deadline is close. For each case, start from the documented gate language. Ask who can expand the canary and which evidence is required to reverse a fail. Any step that depends on an unnamed person becomes a release condition.

Ask the room to replay one historical promotion under the proposed gate. If the historical case would have shipped while later causing a known incident, revise the gate before treating it as policy.

Compare the proposal with the ML ops lead seat and a prompt change review. If the model depends on shared features, review a feature store cutover. More operating decisions live in the war-game decisions hub.

Publish the gate only after a dry run can reproduce the same pass or fail label from the stored artifacts without manual reinterpretation.