Engineering

Stress-test a webhook retry policy before you ship it

Stress-test a webhook retry policy before backoff charts, dead-letter promises, and partner SLAs harden into what every delivery path will quote.

Retry policies fail when backoff invents calm the storm never earned, when poison messages dual-count as success and as parked, when partners still see duplicates the docs swore were impossible, and when ops cannot show who owns a stuck queue. A neat sequence diagram is not evidence.

What belongs on the table

One sentence for why the policy exists, which event classes and partners it covers, who owns dead-letter triage and partner comms, and the rollback trigger if duplicate rate or lag past a named threshold. Attach the last load-test notes, backoff schedule, idempotency keys, and the measured path from fail to dead-letter to human review. If platform, partner ops, and support disagree on which retries are real, stop and reconcile first.

Name the decision you will make if the stress test finds nothing new, and the delay criteria if any event class lacks a named triage owner or a verified idempotency path.

Break cases

  • Storm fiction: retries that look gentle while a single outage multiplies traffic past partner caps.
  • Poison loops: messages that retry forever without a timed park.
  • Duplicate theater: partners that still process the same event twice under load.
  • Ownership blur: dead-letter steps that still say "platform" without a named role.
  • Comms lag: partner alerts that trail the delivery miss partners already felt.

Optional counsel seat if contractual delivery SLAs bind the language. Optional security seat if payloads carry secrets near retry logs.

How to run it

Feed Pingpong the draft retry design, load-test notes, and open risk list. Early passes steelman the design. Later passes attack from platform, partner, support, security, and skeptic seats. End with a pass that turns surviving objections into clearer backoff caps, a timed dead-letter SLA, or a hold. Delete invented "partners already handle retries" claims and dual-counted success rates.

Ask platform and partner seats to price the behavior the published policy will invite. If day-one docs promise at-most-once while the last drill showed silent duplicates, partners will treat the promise as false. Write the intended caps, the idempotency checks, and the language you will refuse, then attack whether trust still holds under that discipline.

When the policy coincides with a new region or a partner cutover, force engineering and partner ops seats to map every claim that still assumes the old endpoint list. Secrets stores, status pages, and support macros count. A policy that looks clean in a RFC while endpoints still accept unsigned retries will fail on the first audit.

Force a day-after narrative: what happens if a partner screenshots duplicate charges, if a poison queue fills overnight, or if an auditor asks for delivery logs you cannot produce. If those stories are stronger than your mitigation plan, fix the package before you ship. Related: stress-test a webhook contract, stress-test a rate limit change, stress-test a status page update, pretend you are the CTO, and the war-game decisions hub. Process: how to run a Pingpong.