Pattern · seen in 2 breakdowns across 2 companies

Retry Budget

Each client gets a small number of retry tokens that build back up slowly, so a failing service gets only a little extra load instead of every client piling on to keep it down.

The mechanism

At its core: when a service is failing, unlimited retries multiply the load on it and can pin it down for good. A retry budget caps the extra load each client can add, so any retry storm is avoided.

LIVE ARTIFACTA CAP ON THE STORMOPEN FULL SCREEN ↗
THE IDEAWhen a shared service starts failing, its clients retry. With no limit, those retries multiply the load: one failed request turns into several retries, and the extra traffic keeps the service overloaded, which causes even more failures and more retries. The load can stay pinned above capacity even after the first problem is gone. A retry budget gives each client a small number of retry tokens and stops it retrying once they run out, so the extra load retries add is capped at a small, known amount and no storm gets going.
WHAT TO TRYStart with the budget wide open and watch the retries bury the service. Then tighten the budget and watch the extra load shrink until the service is healthy again.

Open the retry budget wide and the service drowns; tighten it and the service comes back.

Definition

When a request fails, the client often retries - it sends the same request again, hoping the problem was brief. That is usually fine. But if a shared service starts failing under load, every client retrying at once adds to the overload that is already the problem. That extra load causes more failures, which cause more retries, which add even more load. The result is a retry storm: a flood of retries that holds a struggling service down, and can keep it down even after the first problem has passed.

Spacing retries out with backoff and jitter - waiting longer after each failure and not retrying in perfect sync - helps, but over a long outage it still cannot cap the total load the retries add.

A retry budget puts a hard cap on this. Each client gets a small number of retry tokens, an idea called a token bucket - every retry spends one token, and tokens are added back slowly over time. While the client has tokens, it retries failed requests normally, so brief failures are still hidden from users. When the tokens run out, the client stops retrying, or retries only a small fixed share of requests. The effect is a cap you can work out in advance: no matter how long the service stays down, retries can add only a small, fixed amount of extra load, not an open-ended pile-on.

The strongest version does not leave this to each team. It builds the budget into shared infrastructure that every client already uses - the client library or SDK, the service mesh, or the RPC framework. Then bounded retrying is the default everywhere, not something each team has to remember to add.

This matters when many clients share a dependency whose failures can be made worse by load, and where spacing retries out is not enough on its own to cap the total. If losing a few retries would not hurt and retries cannot overwhelm anything, a budget is more complexity than the problem needs.

When it applies

01Many clients depend on the same service. When that shared service fails because it is overloaded, every client retrying at once adds to the overload that is already the problem.
02You own a shared client layer. An SDK, service mesh, or RPC framework is a place to make bounded retries the default, so every team that uses it is protected without extra work.
03Spacing retries out is not enough on its own. Backoff and jitter stop every client from retrying at the same instant, but over a long outage they still do not cap the total load the retries add.

Tradeoffs

A budget can stop retries that would have been fine. During a long run of brief, harmless failures, the budget may cut retries off even though the service could have taken more, so some requests fail that a retry would have saved.
Each client only caps itself. There is no central control over the whole fleet - the fleet-wide limit is just the sum of every client capping its own retries, so it holds only as long as every client actually has the budget turned on.
Picking the size is a judgment call. Make the budget too small and it never lets clients retry, so it helps no one; make it too large and it never actually caps anything, so it protects no one.

The same move, 2 ways

Every row is a production system that bet on this pattern — the note says how, in that system's own terms.

LinkedIn
LinkedIn Engineering
2023
The server-assisted form of the pattern: Hodor refuses a request before the application code runs, so retrying it on another copy is safe whatever it does. Both caller and server keep a retry budget, an idea from Google's SRE book, to bound the storm risk. When the server's budget runs out, that itself signals a widespread overload and switches retries off entirely, accepting failed requests to protect the traffic still being served. Read the breakdown →
Amazon (AWS)
Amazon Builders' Library
2019
Amazon's local token bucket caps each client's extra load with a real mechanism rather than advice: retry freely while tokens remain, at a fixed rate once they run out. It shipped as default AWS SDK behavior in 2016, so the safe behavior is the default one. The pattern is the difference between telling clients not to storm and making a storm arithmetically impossible for each client. Read the breakdown →

Problems this pattern answers

The walls where its breakdowns live — each opens the cross-company comparison.