Pattern · seen in 2 breakdowns across 2 companies
Retry Budget
Each client gets a small number of retry tokens that build back up slowly, so a failing service gets only a little extra load instead of every client piling on to keep it down.
The mechanism
At its core: when a service is failing, unlimited retries multiply the load on it and can pin it down for good. A retry budget caps the extra load each client can add, so any retry storm is avoided.
Open the retry budget wide and the service drowns; tighten it and the service comes back.
Definition
When a request fails, the client often retries - it sends the same request again, hoping the problem was brief. That is usually fine. But if a shared service starts failing under load, every client retrying at once adds to the overload that is already the problem. That extra load causes more failures, which cause more retries, which add even more load. The result is a retry storm: a flood of retries that holds a struggling service down, and can keep it down even after the first problem has passed.
Spacing retries out with backoff and jitter - waiting longer after each failure and not retrying in perfect sync - helps, but over a long outage it still cannot cap the total load the retries add.
A retry budget puts a hard cap on this. Each client gets a small number of retry tokens, an idea called a token bucket - every retry spends one token, and tokens are added back slowly over time. While the client has tokens, it retries failed requests normally, so brief failures are still hidden from users. When the tokens run out, the client stops retrying, or retries only a small fixed share of requests. The effect is a cap you can work out in advance: no matter how long the service stays down, retries can add only a small, fixed amount of extra load, not an open-ended pile-on.
The strongest version does not leave this to each team. It builds the budget into shared infrastructure that every client already uses - the client library or SDK, the service mesh, or the RPC framework. Then bounded retrying is the default everywhere, not something each team has to remember to add.
This matters when many clients share a dependency whose failures can be made worse by load, and where spacing retries out is not enough on its own to cap the total. If losing a few retries would not hurt and retries cannot overwhelm anything, a budget is more complexity than the problem needs.
When it applies
Tradeoffs
The same move, 2 ways
Every row is a production system that bet on this pattern — the note says how, in that system's own terms.
Problems this pattern answers
The walls where its breakdowns live — each opens the cross-company comparison.