Dead Letter Queue vs Retry Queue: The Real Difference
Dead letter queue explained for carrier integration: how it differs from a retry queue, with a worked webhook example and FAQ.
Ask ten engineers to define a dead letter queue and you'll get ten slightly different answers, most of which also describe a retry queue. That confusion costs real debugging time when a carrier webhook silently stops arriving at 2am. Here's the precise version, and where it splits from a retry queue.
What is a dead letter queue?
A dead letter queue (DLQ) is a separate, terminal queue that holds messages a consumer could not process successfully after exhausting its retry budget, so they can be inspected and replayed by a human rather than lost or retried forever. A dead letter queue (DLQ) is a separate queue that holds messages a consumer could not process successfully, and when a message has been retried up to a configured limit, or has been explicitly rejected, the broker stops redelivering it and moves it off the source message queue to the dead letter queue, instead of discarding it or retrying forever.
The purpose isn't retry logic. It's containment. A dead letter queue (DLQ) is a secondary queue used in messaging systems to store messages that cannot be successfully processed or delivered by the primary system, often referred to as "dead letters" because the main system has given up on them, typically after repeated failures caused by unresponsive consumers, malformed messages, or invalid routing. This is not a bucket for things that will succeed on the next try. It's a bucket for things that need a person, a schema fix, or a manual decision before they move again.
DLQ vs retry queue: the distinction that matters
A retry queue is intermediate and automatic: it expects the message to eventually succeed. A dead letter queue is terminal and requires intervention: the retry budget is gone and nothing will move the message forward except a human or a triage job. A Retry Queue is a mechanism that automatically retries failed message processing with configurable strategies before moving messages to the DLQ, with automatic retry with backoff strategies, configurable retry limits, and handling for transient failures like network issues or temporary service unavailability. A DLQ, by contrast, holds messages that cannot be processed successfully after exhausting all retry attempts, functioning as a "failed messages" inbox that requires manual or automated intervention.
| Property | Retry queue | Dead letter queue |
|---|---|---|
| Lifecycle stage | Intermediate, expects success | Terminal, expects intervention |
| Movement | Automatic, feeds back to main queue | Manual replay or automated triage job |
| Typical trigger | Timeout, 5xx, connection reset | Retry budget exhausted, 4xx malformed, schema mismatch |
| Who acts on it | The broker | An engineer or an alert-driven job |
| Failure type it handles | Transient | Permanent or semi-permanent |
Why carrier integration platforms need both
The split matters because carrier webhooks fail for two structurally different reasons, and treating them the same wastes either time or capacity. Transient failures are a network blip, a downstream service restarting, a database connection exhausted for a few seconds — the message is fine, the environment was briefly not, and these should be retried with backoff and should not reach the DLQ. A carrier sandbox returning a 503 on a Tuesday afternoon is transient. A tracking payload missing a field your consumer expects is not, and no amount of retrying fixes it.
Semi-permanent failures, such as a downstream dependency that has been down for twenty minutes or a reference record that has not been created yet, will eventually succeed on retry but not within any reasonable delivery limit — this is exactly the middle ground where a good retry policy with exponential backoff earns its keep before anything reaches the DLQ.
In a multi-tenant carrier middleware, this distinction also protects tenant isolation. Endpoints that enqueue incoming webhooks for background processing should give that internal queue a DLQ too, because a malformed payload or a bug in the handler will otherwise either poison the queue, blocking everything behind it, or get dropped entirely. One shipper's broken webhook receiver shouldn't be able to stall delivery for every other tenant sharing that queue.
A worked example: a tracking webhook that won't process
Take a PostNL "delivered" tracking webhook arriving at a middleware ingestion endpoint. The signature checks out, but a recent schema change on the consumer side means the handler expects a tenant_id field the payload no longer carries. The handler throws on every attempt.
- Attempt 1: fails immediately, requeued with a short backoff.
- Attempt 2: fails again, backoff doubles.
- Attempt 3: fails again, retry budget exhausted.
- The message moves to the DLQ, not back to the main queue.
What should land in that DLQ record isn't just the payload. On AWS SQS, you attach a redrive policy to your main queue that names a target DLQ and a maxReceiveCount, and when a message's receive count exceeds that threshold, SQS automatically moves it to the dead-letter queue while keeping the original message ID intact. A usable DLQ record for a webhook needs the full context, not just the body:
{
"event_id": "evt_9f2c1a",
"idempotency_key": "postnl-track-9f2c1a",
"destination_url": "https://tenant-142.middleware.eu/webhooks/postnl",
"headers": { "signature": "sha256=...", "timestamp": "..." },
"payload": { "status": "delivered", "shipment_id": "3S..." },
"dead_lettered_at": "2026-09-28T11:04:12Z",
"attempts": [
{ "ts": "11:02:01", "status": 500, "latency_ms": 340 },
{ "ts": "11:02:31", "status": 500, "latency_ms": 288 },
{ "ts": "11:03:31", "status": 500, "latency_ms": 301 }
]
}
Because a message can crash after partially succeeding, replaying it blindly risks double-processing, so before any replay logic runs, your consumer needs to check an idempotency key, almost always the provider's event_id, against a record of what's already been processed. That's the whole reason idempotency keys and DLQs are usually discussed together: one without the other turns a fix into a new incident.
What goes wrong without one
Without a DLQ, a single message that always fails has two possible fates, and both are bad: it is redelivered endlessly, blocking the queue and burning consumer capacity, or it is dropped and the failure disappears silently — a dead letter queue gives that message a third destination, somewhere it can be inspected, diagnosed, and replayed once the underlying problem is fixed.
There's a well-known trap on the other side of the fence too. A poison message that reliably crashes the receiver will fail every attempt, land in the DLQ, get replayed once someone fixes what they think is the bug, and fail again. A poison pill sitting at a given offset causes the consumer to throw on deserialization and die before committing the offset, so the app resumes from the same point, reads the same poison pill, and dies again, flapping in a loop while the offset never advances. Streaming systems hit this hardest, but the same failure shape shows up in any naive auto-replay job for webhooks. Treat DLQ replay as a deliberate, monitored action, not a cron job.
How major brokers implement dead-lettering
Mechanics differ by broker, and the differences aren't cosmetic. Quorum queues in RabbitMQ 3.10 introduced at-least-once dead lettering as an opt-in feature for source quorum queues, ensuring that all messages dead lettered in the source queue will arrive at the target queues eventually, whereas prior versions could lose dead-lettered messages in transit, via RabbitMQ's own documentation on the change. That closed a real gap: before 3.10, a dead-lettered message could vanish between the source queue and the DLQ itself, which defeats the entire point of having one.
On the carrier-facing side, most platforms that fan out webhooks to shippers implement some version of this pattern operationally, even if they don't expose the broker internals. When evaluating carrier integration software such as nShift, EasyPost, ShipEngine, Sendcloud, or Cargoson, expect this as table stakes rather than a differentiator: retry with backoff for transient failures, a DLQ for exhausted or permanently broken deliveries, and visibility into both.
FAQ
Is a dead letter queue the same as a failed-jobs table?
Functionally similar, mechanically different. A DLQ is broker-native (SQS, RabbitMQ, Kafka all ship the primitive). A failed-jobs table is application-level, usually a database row your own code writes when a job handler gives up. Both serve the same purpose: a place for failures that need eyes on them.
How long should messages sit in a DLQ before expiry?
Longer than the source queue's retention, always. Alert when DLQ depth exceeds roughly 10 events, since a small standing backlog is usually the first sign of a quietly failing dependency, and alert when the oldest event has sat unreviewed for more than about an hour. Depth and age, not raw volume spikes, are the metrics worth watching.
Should DLQ replay be automatic or manual?
Manual, or at minimum manually triggered after inspection. Automatic replay of an unfixed poison message just recreates the same failure loop described above.
Does a DLQ guarantee no data loss?
No. It only covers messages that reached the queue in the first place. If a webhook never arrives because of an upstream carrier outage, a DNS failure, or a dropped connection before the message entered your system, there's nothing for the DLQ to catch. That gap needs its own monitoring, typically a synthetic check or reconciliation job against the carrier's own event log.