A timed-out request tells you the response never came back, not whether the work got done.
Cause it · Build it · Survive a day · Compare with five real systems
A request that returns a clear error is easy. You know it failed, so you retry it. The hard case is the request that times out with no error. The connection drops, the server dies partway through, or the response is lost on its way back. All three leave the caller holding the same timeout. The only option left is to retry, and a retry is exactly what can run an operation a second time when it already worked the first. When the operation moves money, guessing wrong either way costs: retry something that already succeeded and you charge the customer twice, give up on something that actually worked and the customer has paid for nothing.
This page follows five companies into the same problem. Stripe and AWS are called by outside developers, so their fix is a rule in the API contract itself: send a key. Airbnb and Shopify sit on both sides (their own services call their own payment code), so they build the machinery inside those services. Segment can't ask its callers, phones and third-party devices that won't follow rules, to do anything, so it catches the duplicates downstream instead. Their answers, with the trade-offs each one makes, are compared below.
The wall
Every version of this problem reduces to one picture: a call that changes something crosses a network, and there are three places it can die. Only the first is safe to retry, and the caller can't tell the three apart, because all three look identical from the outside: a timeout.
Cause it yourself before reading anyone's answer: cut a $100 charge at each of the three points and decide what the client should do. The rescue tool is an idempotency key, a name the caller picks for the operation and sends unchanged on every retry. Reach for it and a second decision appears: where the server stores that name so it can recognize the retry later. Every failure you can cause here is one the five posts describe.
Cause the failure yourself before reading how anyone fixed it.
0.6%
of all events arriving in a four-week window were duplicates the server had already received
Segment · 2017
a few seconds
of replica lag is enough to turn a correct idempotent retry into a double charge
Airbnb · 2019
many times a day
a one-in-a-million failure still happens many times a day at Shopify's payment volume, so the ambiguity is constant, not an edge case
Shopify · 2022
99.999%
payment consistency achieved since the framework launched, while annual payment volume doubled (the payoff side of solving this wall)
Airbnb · 2019
What it costs to ignore, the sharpest way the fix itself breaks, why scale makes it unavoidable, and what solving it earns. Each number is measured by the company that reports it.
Build the defense
You caused the failure from the victim's seat. Now let's design the solution. You need to make six decisions (the box below lists them), and the goal is a day of traffic survived with no double charges. Run it first with the naive defaults, the obvious choices you'd make before learning anything. Then observe what breaks. Each failure is your next lesson: it names the company that hit it, links to their post, and points at the one decision that would have prevented it. Clear the day, and the design faces five attacks, each a real-world failure that one of the five companies actually hit.
Limited by size: drop the oldest keys when full, and alert an engineer if keys start expiring in under a day
Forever
The day, six events
A normal charge
A lost request
A crash mid-charge
A lost reply
Two genuine orders
A late retry
Five attacks, from the posts
Stripe 2017: a reused key
Airbnb 2019: reads moved to a read-only copy
Segment 2017: traffic 10× for a week
AWS 2021: a known key with a different amount
Shopify 2022: a retry after the window
Survive the day and the attacks, and your design becomes the sixth column in the comparison below.
Once your design survives a day, you can stop here. By then you have caused the failure and built a design that survives it. That's the core of it. Everything below is for going deeper: the five real answers side by side (the comparison), which answer fits which constraints, and the original posts. If you are preparing for an interview, the interview note walks the same trades in interview order.
Stuck on a failure card? The five companies below made the same decisions you are making, and wrote down why. Reading their posts while debugging is how most engineers first find them.
Same wall, five answers
All five companies start from the same idea: give each operation a name (the key), and make the system remember it. After that, they answer every hard question differently. Below, in order: who catches the duplicates, the five designs drawn the same way, a table that compares them, and each company's full answers.
Who catches the duplicates?
StripeAWSAirbnbShopifySegment
All on the callerAll on the server
Stripe puts the job on the caller: send one key per charge, and the same key on every retry. AWS asks the same, but its SDK creates the key for you, so most callers do it without knowing. At Airbnb and Shopify the callers are their own services, so most of the machinery sits inside the payment service. Segment's callers are phones that lose signal and apps it doesn't control, so it asks them for almost nothing (a random ID, and even that is optional) and catches the duplicates itself, later in the pipeline. The diagrams below follow the same order: read down and watch the job move from the caller to the server.
Five designs and yours, drawn the same way
Where the key is made Where the key's memory lives What a duplicate gets back What still breaks it
Duplicates are dropped later. The caller is never told.
YounowYour design
A dashed mark means the company's post doesn't say. We leave a gap rather than guess. The table below shows the same comparison at a glance, and the full answers follow it.
Start with the first row. Who the callers are decides almost everything below it.
At a glance
Stripe
AWS
Airbnb
Shopify
Segment
YOU · survive a day first
Who calls
Other companies' developers
Developers and the SDKs they use
Airbnb's own services
Its own services, calling out to a payment partner
Not stated means the post doesn't say, and we don't guess. Columns run from the caller's side on the left to the server's side on the right.
Q1Who names the operation?▸
Only the caller knows what it meant to do, and every post that takes a position agrees. The differences show up when the caller can't, or won't. In the mission, this is IDENTITY.
Stripe2017The caller creates the key and sends it as an Idempotency-Key header on every request that changes something. Keeping keys straight is the caller's job: reuse an old key for a new charge and you silently get the old charge's response back.
AWS2021The caller sends a ClientToken (AWS's name for the key). Two requests with the same token are duplicates by definition. When the caller doesn't send one, the SDK and command-line tool create one and reuse it across automatic retries, so the application never knows a retry happened. AWS says plainly that identical request parameters do not mean identical intent, which is why a token beats a fingerprint built from the parameters.
Airbnb2019The caller creates one unique key per request, and chooses what kind. Request-level keys (random, one per request, reused on every retry) catch retries of that one call. Entity-level keys (built from the thing being acted on, like payment-1234-refund) make sure a given refund happens only once, ever, even across requests that never talked to each other.
Shopify2022The post tracks each attempt under an idempotency key but doesn't say who creates it. What it does say: the key is a ULID, a kind of ID that sorts by time, which keeps the database index compact. That choice alone cut insert time by 50%.
Segment2017The caller's SDK tags every event with a messageId, a plain random ID. It is deliberately the smallest thing Segment could ask of a caller, so anyone in any language can send data. If the caller doesn't send one, Segment's API adds it.
Q2Where does the key's memory live?▸
A good interviewer's next question, because the wrong place can undo the whole guarantee. In the mission, this is MEMORY and READS.
StripeA store on the server that maps each key to how far its request got, checked on every request. How long to keep it is part of what the API promises its callers: the saved response must outlast any realistic retry, or a late retry gets treated as new.
AWSIn the same commit as the change itself. The token and the change are saved together or not at all. The original request's parameters are saved alongside, so a reused token with changed parameters can be caught.
AirbnbThe main database only, never a read-only copy (a replica). A copy runs a few seconds behind, and that delay turns a correct retry into a double charge. So Orpheus, Airbnb's library for this, reads and writes the key's memory on the main database only, and wins back the lost capacity by splitting the key tables across more machines.
ShopifyIn the payment service's own database, recording which steps of the attempt already ran. The ULID format is itself a storage choice: keys that sort by time suit the way the database indexes them.
SegmentA ledger on each worker's own disk (RocksDB), holding about 60 billion recent IDs in 1.5 TB. Messages are routed by ID, so each worker only has to answer "seen this?" for its own share. The output log (a Kafka topic), not the ledger, is the source of truth.
Q3The server crashes halfway through. What does the retry find?▸
The hardest of the three failures you caused at the top of the page. The key tells the server this is a retry, but what the server does next is where the real engineering is. In the mission, this is MEMORY.
StripeNames the problem but doesn't solve it: recovery is "heavily dependent on implementation." If the database rolled back the half-done attempt, run the whole thing again. Otherwise the server must pick up where the half-done attempt stopped and finish it.
AWSDesigned so it can't half-happen: saving the token and making the change are one all-or-nothing step. The retry either finds a finished request and sends back its response, or finds nothing and starts fresh. Either half on its own would break the promise, so the two are never allowed to separate.
AirbnbEach request is built so it can be interrupted safely, in three all-or-nothing phases: before the network call, the call itself, and after it. Each database phase is one transaction, and the rule is simple: no network calls inside a transaction, and no transaction spanning a network call. Interrupt it anywhere and the system can recover. A time-limited lock on the key (a lease) stops two attempts running at once.
ShopifyThe key records which steps ran. A retry with the same key first runs recovery steps that rebuild where things got to, so only one request ever reaches the payment partner.
SegmentAdmits that its three steps (write the ledger, publish the output, confirm the input) can't happen as one all-or-nothing step. So it picks one source of truth: on restart, the worker repairs the ledger from the output log, which is also the permanent record.
Q4What does a duplicate actually get back?▸
Sounds like a detail, but AWS builds its whole argument on it: what the reply looks like decides whether retrying can be the default. In the mission, this is REPLY.
StripeThe saved response from the finished operation, sent again as if it were the first.
AWSA success, every time, meaning the same thing as the first. Even if someone has since shut the instance down, the retry still gets a success, and the body just says "terminated." An "already exists" error would also be idempotent, but it forces every caller's code to handle a special case, which is exactly what makes retry-by-default hard to offer.
AirbnbThe saved final response, handed back to any later retry. The downside: a table that grows with traffic forever and is hard to trim or reshape later.
ShopifyNot stated, beyond the guarantee itself: however many retries arrive, only one request reaches the payment partner.
SegmentNothing. The duplicate is dropped later in the pipeline and the caller never learns it sent one. Its callers couldn't do anything with that information anyway.
Q5How long does the protection last?▸
Every key's memory runs out, on a clock nobody sees. Five different answers to when, and only one of them is a fixed number. In the mission, this is WINDOW.
Each bar shows how the window ends, not how long it is. The bars share no scale, because only Shopify's window is a number.
StripeNo number, but a clear rule: keep keys longer than any realistic retry, and treat how long you keep them as part of what the API promises, not an operations afterthought.
AWSAs long as the resource exists, plus a margin, after which any late retry would either have arrived already or no longer make sense. Keep tokens too briefly and a late retry creates a duplicate. Keep them forever and a new token can clash with an ancient one.
AirbnbTwo clocks. A lease (a time-limited lock) on each key while its request is running: it outlasts the network call's timeout, then expires, so a crashed server can't hold the key forever. And a cap on total retry time, so retrying can't go on forever.
ShopifyAbout 24 hours, usually. The post treats the length as a real dial: longer protection, traded against the expense of tracking every attempt. The window is a design decision, not an accident.
SegmentSet by size, not time. Each worker caps its ledger and deletes the oldest IDs first, so a spike in traffic shrinks the window instead of toppling the system. An alert fires if the window drops under 24 hours.
Q6What still breaks it?▸
The row interviewers spend the most time on. Each company's own post names the failure its guarantee doesn't cover. Knowing these is what separates using the pattern from understanding it. The red marks in the diagrams above show where each one sits. In the mission, these are the five attacks.
StripeCareless keys on the caller's side. Reuse an old key for a new charge and the server silently sends back the old charge's response. The server can't detect this. It depends on every developer who calls the API using keys properly.
AWSA reused token with changed parameters. The safest reading is that the customer meant something different, so the service refuses with a validation error that names the mismatch. The guarantee protects intent, and changed parameters are a new intent.
AirbnbWhere the key's memory lives. Keep it on a copy that runs behind and the very failure the key was meant to cure comes back, so reads go to the main database only. And a mislabeled error: mark a failure as safe to retry when it isn't, and double payments are back, with manual cleanup behind them.
ShopifyTime. A retry that arrives after the window of about 24 hours simply isn't covered. Shopify's short timeouts cause more retries on purpose, which is only safe while the keys are still remembered.
SegmentPressure on the window. "Almost" exactly once is the honest name. Under heavy traffic the window shrinks by design, and the 24-hour minimum is watched by an alert, not guaranteed.
Which answer is yours
The five designs aren't competing opinions. Each is a spot on the caller-to-server line above, and each company's callers put it there. Find the row that matches your situation.
You publish an API that other people's code calls, and you can ask callers to cooperate
Stripe for the basic rule, then AWS for the two additions that make it easy to live with: the SDK fills in the key, and a duplicate always gets a success back, so callers never need special-case code.
The callers are your own services, and the data underneath is split across machines and copied
Anywhere a call that changes something can time out: reserving stock (the last item sells twice), sending an email (it arrives twice), delivering a webhook (now it's the receiver's problem), starting a server (AWS's own case), recording a payment in the books. Ask the six questions. The answers change. The wall doesn't.
What to steal
Five rules from these posts that carry over to any system you build. Each one is a row of the table, boiled down.
Let the caller name the operation. Don't compute it. A fingerprint built from the request's details can't tell an accidental repeat from a customer who really wants two identical orders. Only the caller knows what it meant, so the caller (or its SDK) names the operation. Q1
Save the key's memory in the same commit as the change, or plan for the moment they come apart. AWS makes them one all-or-nothing step. Airbnb splits the request into three all-or-nothing phases. Segment repairs its ledger from the output log after a crash. Nobody leaves the gap unhandled. Q3
The guarantee is only as good as the place the key's memory lives. Keep it on a copy that runs behind and the failure it was built to cure comes back: a few seconds of delay is a double charge. Always say where the memory lives. Q2
Decide what a duplicate hears. The saved response, a refusal, or silence can each be right. But what the reply looks like decides whether every caller can safely retry by default. It's an API decision, not an accident. Q4
Every key's memory runs out. Protection ends, whether by policy, a lock timer, size, or about 24 hours. Pick the window on purpose, write it down, and set an alert for when it shrinks. Q5
If this comes up in an interview
The question, as asked
"How do you prevent double payments?"
"The client's request timed out. Did the charge happen?"
"Design an idempotent API for creating orders."
The 90-second shape
The caller names the operation and sends the same name on every retry. The server remembers the name, and I'd say where: in the same commit as the charge, or in a separate store with recovery steps, always read from the main database. A retry gets the saved response. The memory has a window, so I'd say how long, and what happens after it ends. Then I'd say what still gets through: a reused key, a read from a read-only copy, a changed parameter, a retry after the window.
The follow-ups are the attacks
Every attack you held above is a question a good interviewer asks after "use idempotency keys."
The interviewer asks
Attack
What held
"Same key, new charge. What happens?"
1 · Stripe
— Nothing can stop it. Using keys properly is the caller's job.
"Reads move to a read-only copy to take load off the main database"
2 · Airbnb
— Reads stay on the main database
"Traffic goes 10× for a week"
3 · Segment
— A window set by size, with its cost on the bill. Or forever, with its cost in red.
"Same key, different amount"
4 · AWS
— Refuse with a validation error
"A retry arrives after the window"
5 · Shopify
— Reconciliation, which catches it afterwards but can't prevent it
Senior vs Staff
Senior: "I'd use idempotency keys." Staff: "Keys, saved in the same commit as the charge, read from the main database, the saved response sent back on a retry, a 24-hour window, and reconciliation after it. And here's what that costs: the charge in the same database as its record, responses stored for every request, and a reconciliation job that never ends." The difference is the bill. A Staff answer comes with one. The bill in the mission above is yours.
Answers that sound right
"Hash the request to detect duplicates." → two identical real orders become one.
"Read the key from a replica. It's only a read." → a few seconds of delay is a double charge (Airbnb, 2019).
"Keep keys forever so it can never happen." → storage that grows forever, and clashes years later (AWS, 2021).
"Return an error on the retry. That's still idempotent." → half right. It is idempotent, but now every caller needs special-case code. A trade-off, not a mistake (AWS's own argument).
Patterns in this class
The reusable building blocks behind these five designs. Idempotency keys are the core. The rest cover the failures keys alone let through.
Retry a job that should run once into two servers, break the all-or-nothing link between the token and the change, and watch a fingerprint of the request swallow a customer's genuine second request.