The Selfish Retry: Timeouts, Backoff, and Jitter at Amazon
Marc Brooker's Builders' Library article is Amazon's playbook for the three tools every network call needs. Timeouts, so a failure shows up instead of hanging forever. Retries, so brief random failures get hidden. And backoff with jitter, so that retrying does not set off a thundering herd and cause the next outage. Its sharpest idea is that retries are selfish. A client that retries spends more of the server's capacity to improve its own odds, which piles on extra load exactly when the server can least handle it. Amazon's answers are: - Retry only when the dependency looks healthy and stop when retrying is not helping. - Cap each client's retries with a local budget (a 'token bucket') built into the AWS SDK since 2016. - Treat any call with side effects as unsafe to retry unless it is idempotent (the way EC2's RunInstances is) using client-generated idempotency tokens. - Add random jitter (small random offsets) to every timer and scheduled job (not just retries) because clients on the same schedule pile up together.
Brown out a dependency under a retrying client fleet, then switch on backoff, jitter, and the token bucket one at a time, and watch the retry storm dissolve.
Problem
Whenever one system calls another, a failure can come from anywhere: the servers, the network, the load balancers, the software, the operating system, even the people operating it. Amazon's stance is that you cannot build systems that never fail, so systems must tolerate failure without letting a small number of failures grow into a full outage. The hard part is that this growth is caused by the very tools meant to tolerate failure.
Many failures first show up as slowness: requests taking longer than usual, maybe never finishing. A client waiting on such a request ties up its own resources (memory, threads, connections) the whole time, and a server that keeps waiting piles up work for callers that have already given up and moved on. Waiting without limit turns one component's slowness into a global resource shortage, which is what proper timeouts prevent. But timeouts create an ambiguity that retries then have to deal with: a call that timed out may or may not have actually run, and its effects may or may not have happened.
Retries handle the availability side of that problem extremely well: repeat the call and most brief failures disappear. The cost is the catch: a retry makes the server do the work again, spending capacity that every other client also needs. Each retry asks an already-struggling server to spend its capacity twice, or three times, for one client's benefit. If the real failure is overload, a fleet of retrying clients turns a partial slowdown (a 'brownout') into a lasting storm, and the system can stay down under the weight of retries long after the original cause is gone. Backoff, spacing retries out with waits that grow longer each time, keeps the load on the server more even, but it has its own failure mode: every retry lining up at the same time. Clients that failed at the same moment back off on the same schedule and come back at the same moment, re-creating the pile-up in repeated waves. And Amazon traffic is not smooth in the first place. Normal client behavior, recovery after a failure, and even ordinary scheduled jobs all produce big bursts, so anything that lines clients up further is building the next spike.
Solution
Timeouts come first, and Amazon has a strict rule: put a timeout on every remote call, and even on calls between processes on the same machine, covering both connecting and waiting for the reply. The timeout value is chosen from real observation. You look at how long the dependency's calls actually take, and pick a cutoff where only a small, chosen fraction of genuinely healthy requests would be cut off. Set it too generous and resources pile up behind slow calls. Set it too aggressive and the timeout itself creates retries.
Retries come next, and they are applied with great care. Because retries add load to a failing dependency, Amazon retries only when there is evidence the dependency is healthy, and stops retrying when the retries are not improving availability. Spending capacity on retries that cannot succeed only deepens an overload. The rule is enforced by a proper mechanism: each client limits its own retries with a local 'token bucket.' It retries freely while it has tokens, and at a fixed, capped rate once they run out. That behavior shipped inside the AWS SDK in 2016, so every customer using the SDK gets the protection by default, which makes the safe behavior the default one.
Whether it is even safe to retry is more a question for the API than about the client. A timeout does not mean the effects did not happen, so any call that changes something is unsafe to retry unless it is idempotent: guaranteed to take effect only once, no matter how many attempts arrive. Read-only calls are idempotent by default. Calls that create things (or have other side effects) may not be. That is why EC2's RunInstances offers a token: the client supplies a token, every retry carries the same token, and the second launch request is recognized as a repeat and ignored, instead of starting a second fleet of servers.
Backoff spreads retries out over time, with waits that grow but are capped, which keeps the load on the server more even. But backoff alone leaves the lining-up problem: clients that failed together still come back together. Jitter, a small random amount added to each wait, spreads that synchronized wave out into a roughly even, steady stream of retries, and Amazon has a separate analysis of how much jitter to add and how to add it. Then comes the idea that lifts the piece beyond basic retry hygiene: jitter is not only for retries. Traffic to Amazon's services arrives in bursts short enough to hide inside averaged-out metrics, and a lot of that burstiness is caused by the clients themselves. Clients that call on a regular interval line up with each other, and scheduled jobs bunch up on the hour and in the first seconds after midnight. On systems like EBS and Lambda, deliberately jittering those periodic jobs let the same work finish using less server capacity. So Amazon adds jitter to all timers, scheduled jobs, and delayed work. The jitter is chosen the same way every time for a given machine rather than freshly random, so that under overload a human can still make sense of the pattern.
Tradeoffs
- Every timeout will sometimes cut off a healthy request, so choosing one means accepting a rate of false failures. Choosing the cutoff from real timing data means deliberately deciding what fraction of healthy-but-slow requests will be killed and retried. That is a cost paid in repeated work and, for calls that are not safe to repeat, in not knowing whether the first attempt already happened. Generous timeouts cost differently: resources blocked behind slow calls, and failures detected so late that it barely counts as detection. Amazon's contribution is refusing to let the number be a guess. The trade-off itself cannot be avoided.
- Only retrying when the dependency looks healthy gives up some power to hide failures in exchange for stability. Retrying only when the dependency looks healthy, and stopping when retries are not improving availability, means deliberately not hiding the very failures customers feel most, the deep ones caused by overload. That is the point. During an overload, the retry that might have saved one request instead taxes every request. The client accepts worse odds per request in the bad times in order to keep the bad times short. That choice only makes sense across the whole fleet, which is why it lives in the SDK rather than in each team's judgment.
- The token bucket turns retry capacity into a budget, and a budget behaves awkwardly at its limits. While tokens last, retries are free and brief failures get fully hidden. Once the tokens run out, retrying continues at a fixed rate. That caps the extra load during a real overload, but it also throttles retries during a long run of genuinely brief failures, where more retrying would have been safe. And each client's budget is local: each client limits only its own retries without seeing the fleet's total, so the limit on the combined load is hoped for rather than guaranteed.
- Idempotency moves the retry problem into the API's contract. Token mechanisms like RunInstances' make retries idempotent by making the server remember state, an expiry policy, and rules that every new operation now has to define. Saying 'reads are safe, and creates use a token' sounds clean until a call has side effects that are easy to miss, like recording usage for billing, sending notifications, or writing audit logs. At that point 'safe to retry' becomes something you have to prove for each endpoint, not a property you get for free. Amazon is honest about this: good API design is a precondition for safe clients, not a substitute for them.
- Jitter makes the fleet's load smoother but makes each machine's timing less predictable, which is why Amazon insists the jitter be consistent per machine. Randomizing every timer and scheduled job flattens the fleet's combined load, but a system where every action is nudged by a random amount is harder for a human to reason about during an incident. Choosing the jitter the same way every time for a given machine keeps each machine's behavior readable (the same host always fires at the same offset) while the fleet as a whole stays spread out. It is a small design detail carrying a big operating principle: an optimization must not cost you the ability to debug the system it optimizes.
- All three tools manage one trade-off that none of them removes: hiding failures versus revealing them. Timeouts reveal slowness that patience would hide. Retries hide failures that revealing would let you fix. Backoff and jitter hide the cost of that hiding. A system tuned for maximum hiding looks perfectly healthy right up until it collapses, and one tuned for maximum revealing keeps paging humans over noise. Amazon's quiet position, retry but only while it clearly improves availability, is a rule for continuously re-deciding where on that line to stand, not a way to settle it.
Patterns in this article
- Retry with Backoff and Jitter
This is the pattern written up as standard practice by the company that popularized it: capped exponential backoff to even out load, and jitter to break the lining-up that makes backed-off clients return in synchronized waves. Its extra insight is that jitter belongs on all periodic or asynchronous work, not just retries. The same pattern shows up on the API-provider side too. Here it is the job of the clients themselves.
- Idempotency Keys
EC2 RunInstances' client token is the same mechanism Stripe exposes as its Idempotency-Key header. The client names the operation, every retry repeats the name, and the server ignores the duplicates, turning a call that was unsafe to retry into a safe one. Same contract, two companies: retry safety is designed in at API-design time, not bolted on by the client.
- Retry Budget
Amazon's local token bucket caps each client's extra load with a real mechanism rather than advice: retry freely while tokens remain, at a fixed rate once they run out. It shipped as default AWS SDK behavior in 2016, so the safe behavior is the default one. The pattern is the difference between telling clients not to storm and making a storm arithmetically impossible for each client.