Don't blindly retry a send: idempotency for side effects
— Backend, Reliability, AI Agents — 7 min read
Retries are the first thing we reach for when a network call fails. For reads, that's fine. Reading the same row twice hurts nobody.
For anything that changes the outside world, a retry can repeat the change. Send an email, charge a card, post a message, create a ticket: if the first attempt timed out, you don't know whether it happened. Retry blindly and the customer gets two receipts, or two charges.
Idempotency keys, the outbox pattern and an explicit "unknown" state make retries safe. The same rules apply to LLM agents calling tools, where they're easy to forget.
A timeout is not a failure
When a request times out, three things could have happened:
- The request never reached the server.
- The server did the work, but the response was lost.
- The server is still doing the work.
From the client's side these look identical. All you know is that you didn't hear back.
So treat a timeout as unknown, not failed. Errors that clearly happened before the server acted, such as a refused connection or a validation error, are different. Those are safe to retry or report.
Exactly-once delivery isn't on offer
Across a network you get one of two guarantees:
- At-most-once: send once, never retry. Some messages are lost.
- At-least-once: retry until confirmed. Some messages are duplicated.
"Exactly once" doesn't exist at the delivery level (the Two Generals problem is the classic explanation). What systems actually offer is at-least-once delivery plus idempotent processing. The message may arrive twice, but the receiver recognises the second copy and does nothing. That's the "exactly once" in marketing material, and it only works if the receiver dedupes.
Idempotency keys
An idempotency key is a unique value the client sends with a request. The server records it alongside the result. If the same key arrives again, the server returns the stored result instead of doing the work twice.
Stripe's API is the best-known example. You send an Idempotency-Key header, and a retry with the same key returns the original response. The IETF draft on the Idempotency-Key header describes the same pattern for HTTP APIs in general.
Where the key comes from matters more than its format. Derive it from the business intent: order-42:receipt is good, while a fresh UUID per attempt defeats the purpose. Create it before the first attempt and store it, because a key that lives only in memory is gone after a restart. And bind it to the request body. If a client reuses a key with different parameters, reject the request instead of returning a stale result for something else.
If you build the server side yourself, a table and an INSERT ... ON CONFLICT cover most of it:
CREATE TABLE idempotency_keys (
key text PRIMARY KEY,
request_hash text NOT NULL,
status text NOT NULL DEFAULT 'started', -- started | succeeded | failed
response jsonb,
created_at timestamptz NOT NULL DEFAULT now()
);
INSERT INTO idempotency_keys (key, request_hash)
VALUES ($1, $2)
ON CONFLICT (key) DO NOTHING
RETURNING key;One row back means this is the first time: do the work, then store the response. Zero rows means you've seen the key. Load the row, check the hash, and return the stored response, or a "still in progress" error if the status is started.
The outbox: don't send inside your transaction
A common bug: the code commits an order, then sends the confirmation email. If the process dies between the two steps, the order exists and the email never goes. Flip the order and you can send an email for an order that then fails to commit.
The transactional outbox fixes this. Instead of sending, you write a row describing the send, in the same transaction as the business change. A separate relay reads the outbox and does the actual delivery.
BEGIN;
INSERT INTO orders (id, email) VALUES (42, 'a@example.com');
INSERT INTO outbox (idempotency_key, kind, payload)
VALUES ('order-42:receipt', 'email', '{"to": "a@example.com", "template": "receipt"}');
COMMIT;Either both rows exist or neither does. The relay is a normal job worker that retries until the send is confirmed, which makes delivery at-least-once. The idempotency key stored in the outbox row makes those retries safe.
Handling the unknown case
The relay is where timeouts get handled properly:
async function deliver(row: OutboxRow) {
try {
await provider.send(row.payload, { idempotencyKey: row.idempotency_key })
await markSent(row.id)
} catch (err) {
if (isTimeout(err)) {
// It may have gone out. Never send a fresh copy.
await markUnknown(row.id)
} else if (isRetryable(err)) {
await scheduleRetry(row.id, row.attempts)
} else {
await markFailed(row.id, err)
}
}
}What you do with unknown depends on the receiver:
- If it honours idempotency keys, retry with the same key. The worst case is a deduplicated no-op.
- If it doesn't, but you can look things up, ask first ("does a message with reference
order-42:receiptexist?") and send only if it doesn't. - If you can do neither, pick which mistake is cheaper. A missed marketing email is cheaper than a duplicate. A missed password reset is worse than a duplicate. Make that choice per message type and write it down, instead of leaving it to whatever the retry library defaults to.
The same problem in LLM agents
Agents make all of this more likely. An agent loop calls tools, and tools send emails, create records and post messages. Retries come from several places at once: the HTTP client, the agent framework, a workflow engine resuming a crashed run, and the model itself deciding to "try again" after seeing an error.
The model is the worst of these, because it can't tell a timeout from a failure. Show it Error: request timed out and it will often call the tool again.
What has worked for me is moving the safety out of the model and into the harness:
- Mark tools that have side effects. Read-only tools can retry freely. Tools that write go through a ledger.
- Have the harness derive the key, not the model. Use the run id plus the step id, so a resumed run maps to the same key. For tools where the model might rephrase the same intent, also hash the important arguments (recipient, amount, subject) into the key.
- Return "unknown" as its own result. Tell the model plainly: "This may have succeeded. Do not call it again. Check status with
get_message_statusor ask the user." Models follow that instruction far better than they infer it from a timeout. - Put a human confirmation step in front of anything you can't undo, like payments and messages to customers.
Takeaway
- A timeout means "unknown", not "failed". Don't let a generic retry wrapper treat them the same.
- Exactly-once delivery isn't real. At-least-once delivery plus idempotent processing is what you can build.
- Derive idempotency keys from business intent, store them before the first attempt, and reuse them on every retry.
- Use an outbox so the business change and the decision to send commit together.
- In agent systems, the harness owns idempotency. The model never chooses whether a side effect runs twice.