Mannvit retry-safety record

Make the repeat decision visible before a retry exists.

This is Mannvit guidance, not a vendor-specific contract. Complete it for every operation where a repeat could change money, customer communication, ownership, or a consequential record.

DecisionCompleted example: create an onboarding recordCopyable retry-safety record
Stable operation identityThe CRM status-change event ID follows every delivery attempt.<record the event, request, or natural business key>
Protected side effectOne onboarding workspace record may exist for the account.<state what must not happen twice>
Repeat behaviourA completed event returns the stored result; it does not create another record.<state what a duplicate returns or skips>
Unsafe inputMissing onboarding scope is held for review, never written as a blank.<state which payloads must stop rather than overwrite>
Completion proofThe receiver is read back by account ID and linked to the source event.<state the destination evidence that proves completion>
Recovery ownerRevenue operations reviews held records and reconciliation differences.<name the role that owns a non-automatic repair>

Why it matters in practice

Retries are not the exception. They are how reliable integrations work.

A system that retries on failure is more reliable than one that gives up, but only if the operation it retries is idempotent. Otherwise the mechanism designed to recover from failure becomes a source of it.

Why it runs twice

At-least-once is the normal guarantee

Networks time out after the work was done but before the acknowledgement arrived. Webhook providers re-send when they do not get a fast 200. Someone re-runs a batch that half-finished. None of these are bugs. They are how delivery is supposed to work. The message arriving twice is expected, not exceptional.

What breaks without it

The retry becomes a second failure

A non-idempotent write does its work again on every attempt: two invoices for one order, two contacts for one signup, a card charged twice. The damage is worse than the original failure because it is silent, the integration reports success both times, and the duplicate surfaces later as a data problem nobody can trace.

How to make it safe

Give the operation a stable identity

Derive an idempotency key from the event itself, an order id, an event id, not a fresh value per attempt. Record it when the operation completes. On a repeat, the system sees the key exists and returns the first result instead of acting again. For record writes, a natural unique key plus an upsert does the same job.

The retry boundary

A retry is a second attempt at the same operation, not permission to make a second business decision.

In the completed example, a lost acknowledgement may cause the sender to try again. The receiver first checks the stable event identity, returns the prior result if it already completed, and stops work that cannot be applied safely. That is the recovery boundary that prevents a transport failure becoming a commercial duplicate.

Two subtler ways it still goes wrong

Not running twice is necessary, not sufficient.

An idempotency key stops the same operation from acting twice. Two related mistakes survive it, and both corrupt data as quietly as a double.

The harmful no-op. A retry that carries a partial payload can overwrite good remote fields with blanks. An update is only safe to repeat if "no new value" means "leave it alone", not "set it to empty". Skip the write that would blank a field rather than sending it.

The doubling append. A key protects a create or an update, but a plain append, a log line, a ledger entry, an audit row, doubles on every re-run because nothing dedupes it. Stamp each appended row with the event id and refuse a second row that already carries it.

The order of a multi-step apply. If a repeated operation reads and then writes, the read has to come before the destructive step, or a retry acts on stale state. Fix the order once, in code, so a re-run cannot land the steps out of sequence.

The reliability rule

At-least-once delivery can run an event more than once; design the side effect accordingly.

The question to ask of any operation that moves data or money is simple: what happens if this runs twice? If the answer is "a duplicate", the operation needs an idempotency key before it goes near production, not after the first double-charge.