Data operations

War-game a pipeline SLA before teams depend on it

War-game a pipeline SLA by testing each promise against observed delivery times, dependency failures, and the people available to recover the pipeline.

An SLA can sound precise while hiding the clock that matters. “Available by 8:00” may refer to ingestion completion even though finance waits for validation until 10:30. A 99 percent target may also exclude the late partitions that cause the most expensive corrections. The review needs one definition for arrival, one measurement boundary, and a consequence when the service misses.

Freeze the proposal

Write the covered datasets, service hours, freshness target, completeness threshold, incident channel, and escalation owner. Attach thirty days of arrival distributions, the dependency map, recent misses, and actual recovery durations. Identify exclusions in plain language. If a source system has no service commitment, show how its failure affects the promise instead of burying that risk in a footnote.

Name the approval choice and the conditions that force a hold. This prevents the session from drifting into a general debate about data quality.

Seat the contract from both sides

Producer
Defends what the pipeline can operate during normal staffing and peak load.
Consumer
Shows which decisions become unsafe when freshness or completeness slips.
Platform owner
Challenges monitoring coverage, retries, queue limits, and shared dependencies.
Finance or risk owner
Prices a miss and tests whether the proposed remedy has practical value.
Skeptic
Finds the strongest promise with the weakest measurement path.

Run failure cases on the clock

Use Pingpong to walk through a source delay, a partial partition, a bad transformation, and an unavailable owner. For each case, start the SLA clock at the documented event. Ask when the consumer learns about the miss and which message they receive. Then trace the repair, validation, and backfill. Any step that depends on an unnamed person becomes a release condition.

Compare the proposed numbers with a schema migration stress test and the data engineer seat. If recovery depends on infrastructure headroom, review the capacity plan. More operating decisions live in the war-game decisions hub.

Publish the SLA only after a dry run can produce the same breach timestamp, owner, and consumer notice from the monitoring system without manual interpretation.