Retries, backoff and the poison message
At-least-once delivery means your job will run twice. Exponential backoff with jitter, a retry ceiling, and somewhere for the ones that will never succeed.
Almost every queue you will use guarantees at-least-once delivery, because the alternative is losing messages. A worker that crashes after doing the work but before acknowledging it will see that message again, and so will the next worker if the network dropped the acknowledgement. Every job handler therefore has to be safe to run twice: the same idempotency argument as the HTTP lesson, arriving from a different direction.
Backoff should be exponential and jittered. Exponential because a service that is down needs time, and hammering it every second delays its recovery. Jittered because without randomness every failed job retries at the same instant, and you have built a load generator that fires precisely when the target is weakest. A ceiling matters too: after five or six attempts, stop and move the message to a dead-letter queue.
The dead-letter queue is only useful if somebody looks at it. Alert on depth, not on individual failures, and make the payload inspectable: a poison message with a body you cannot read is a mystery, and mysteries do not get fixed. Most teams learn this by discovering four thousand messages in a queue nobody had a dashboard for.
You should now be able to
- Explain at-least-once versus at-most-once delivery
- Design a job that is safe to run twice
- Configure backoff that does not amplify an outage
Loading…