Pattern · seen in 3 breakdowns across 3 companies

Queue with Guaranteed Delivery

A queue with guaranteed delivery saves every message until it has been processed and confirmed, so it survives backlogs and crashes without dropping anything - unlike a buffer, which loses messages under pressure.

The mechanism

The pattern at its core: a producer sending faster than the consumer can keep up, and the difference between a buffer that drops the overflow and a queue that holds all of it.

LIVE ARTIFACTTHE PATTERN, GENERALOPEN FULL SCREEN ↗
THE IDEAWhen work arrives faster than it is processed, something has to give. A buffer has a fixed size, so once it is full the extra messages are dropped and gone for good. A durable queue saves every message, so the same surge just builds a backlog that the consumer drains once it catches up. The failure turns from lost data into delay - which is almost always the better failure.
WHAT TO TRYRun the same producer surge as a buffer, then as a durable queue. The buffer overflows its cap and drops messages for good; the queue holds the whole backlog and drains it later, losing nothing. Same overload, very different outcome.

Send a producer surge through a buffer, then a durable queue - one drops messages, the other just backlogs.

Definition

A queue with guaranteed delivery keeps each message safely stored until it has been picked up, processed, and confirmed done. It can hold a large backlog, keep working when the consumers crash, and never drop a message even under heavy, sustained load. It makes two promises:

  • to producers (the code putting messages in) - once the queue says it has your message, that message will be delivered, no matter what happens next
  • to consumers (the workers reading messages out) - the message stays available, and may be handed to you more than once, until you have processed it and confirmed it done
STORED UNTIL ACKED
A message flows from producer to a durable queue to a consumer, deleted only after the consumer acknowledges it.
A message is written to durable storage the moment it arrives and stays there until the consumer confirms it is done. If the consumer crashes before that confirmation, the message is simply delivered again - it is never dropped.

The pattern is best understood by what it is not: a buffer. A buffer keeps messages in memory (or barely saves them) and acts like a queue on a normal day, but starts silently losing messages once the pressure is sustained. A fast in-memory store used as a queue is the classic example: it works fine as a buffer, but because it runs out of memory and throws old entries away, messages vanish once the backlog grows large enough. Buffers are useful for plenty of things, but they are not real queues, because they do not guarantee delivery.

The principle is simple but easy to break: if your queue's response to pressure is to lose data, it is not a queue, it is a buffer. Real queues save every message, however they are built:

  • a disk-backed log copied across several machines
  • a managed cloud queue you don't run yourself
  • a carefully-run message broker with durable storage

Whichever form you pick, it has to store messages solidly enough to survive consumer crashes, sudden bursts from producers, and part of the cluster going down.

The practical takeaway: any pipeline where a lost message is a problem the customer would notice needs a guaranteed-delivery queue between the producers and consumers. A few examples:

  • search indexing - a lost message means data no one can find
  • notifications - a lost message means an alert that never arrives
  • event-sourced systems - a lost message means the state quietly drifts out of sync

More broadly, any pipeline that does work the system can't recreate from its upstream sources needs one too.

When it applies

01A lost message leaves a visible gap downstream. Indexing pipelines feeding search or analytics, where a dropped message means data that simply never shows up.
02The message is the only record of the work. In async workflows where you can't rebuild the task from anywhere else, losing the message means the work is just gone.
03A dropped message becomes a promise you break. Notifications, alerts, and webhooks that the system already told the user it would send, so losing one is a visible failure.
04Events between services drive real state changes. When one service's events cause another to update its data, a lost event leaves the two out of sync.
05Producers sometimes send faster than consumers can keep up. When work arrives faster than it's processed, you want the queue to absorb the surge, not throw the extra away.
06You're replacing a buffer that already lost data. Swapping out an in-memory queue that has dropped messages under load is one of the most common upgrades a growing system makes.

Tradeoffs

Guaranteed delivery usually costs speed. The queue has to save each message solidly - to disk, or across several machines - before it can tell the producer it has it, which adds delay. That is fine for most workloads, but a very latency-sensitive one might need a different choice.
Not dropping messages just changes the failure: instead of losing them, the queue piles them up. A consumer outage can build a backlog that takes hours to work through, so you have to design for it - scale consumers up under load, watch how deep the queue is getting, and recover cleanly from big backlogs.
Storage cost grows with the backlog. A queue sitting on ten million messages uses real storage - which a managed service bills you for, and a self-hosted one burns disk on. A long incident that builds a huge backlog can cost more than you would expect.
Delivered once is a different problem from delivered at all. Most guaranteed-delivery queues promise at-least-once, meaning a message can arrive more than once, so your consumers have to be safe to run twice (or you accept some duplicate work). Getting true exactly-once takes extra design on top of the queue.
It is real work to run. These queues need durable storage, copies across machines, and careful operational care. The simplest path is a managed service - you pay money to skip the operational burden; running it yourself costs less but hands you the control and the work.

The same move, 3 ways

Every row is a production system that bet on this pattern — the note says how, in that system's own terms.

Discord
Discord Engineering
2025
The Redis-to-PubSub migration. Discord's original Redis indexing queue dropped messages under sustained pressure: a 'buffer', not a real 'queue'. PubSub guarantees delivery and holds large backlogs, so downstream failures now show up as slowdowns rather than lost data. The principle: if your queue's failure mode is losing data, you don't have a queue, you have a buffer. Real queues persist. Read the breakdown →
Meta
Engineering at Meta
2021
FOQS puts this whole property in one platform: items persist, leases turn consumer crashes into redeliveries, ack/nack give at-least-once semantics, and backlogs of hundreds of billions of items are cited as the system absorbing widespread downstream failure. Discord reached for the same guaranteed-delivery property in its PubSub service after Redis dropped messages under pressure; here Meta builds the property itself, as a shared multitenant platform. Read the breakdown →
Slack
Slack Engineering
2017
Delivery is guaranteed by never losing a job silently: the relay advances Kafka's position only after a job is safely in Redis, and re-enqueues a failed job instead of dropping it, so loss turns into delay. Two other articles here (Discord and Meta) make the same trade. The honest exception is at the front door, where the gateway acknowledges a write as soon as the lead Kafka broker has it, accepting a small, documented chance of loss in exchange for speed. Read the breakdown →

Problems this pattern answers

The walls where its breakdowns live — each opens the cross-company comparison.