Pattern · seen in 5 breakdowns across 5 companies

Feedback-Controlled Load Management

A fixed limit on how much load to accept is wrong as capacity varies, so a feedback loop watches a live signal and adjusts the load limit dynamically.

The mechanism

At its core: a fixed limit cannot match a capacity that keeps moving, so it is wrong most of the time. A feedback loop watches the system and slides the limit up and down to follow capacity, keeping the accepted load close to what the system can actually handle.

LIVE ARTIFACTA LIMIT THAT TRACKS CAPACITYOPEN FULL SCREEN ↗
THE IDEAA fixed limit on how much work to accept is a guess about capacity. But real capacity moves - the hardware, the traffic, a slow downstream - so a fixed number is too low half the time (turning away work it could serve) and too high the other half (melting under load). A feedback loop makes the number dynamic: it watches a live signal like queue wait or latency, compares it to a target, and nudges the limit up when there is room and down when there is strain. Steering by the trend keeps it smooth, and because it measures the real system it keeps up with change on its own.
WHAT TO TRYLeave the limit static and watch it sit flat while capacity rises and falls beneath it - too low when capacity is high, too high when capacity drops. Then switch to the feedback loop and watch the limit start to match the capacity line.

A fixed limit sits flat while capacity moves; a feedback loop makes the limit chase it.

Definition

A fixed limit on how much work to accept is an unreliable guess about capacity. Set it too low and the system turns away work it could have handled; set it too high and it melts before the limit ever trips. And capacity keeps moving - faster or slower hardware, a heavier traffic mix, a slow downstream - so any single number is wrong most of the time. Worse, a fixed limit tends to fail all at once: the system keeps accepting work until it suddenly rejects all of it, and the flood of rejections that follows sets off retry storms that recreate the overload.

Feedback-controlled load management replaces the fixed number with a control loop, the way a thermostat replaces a fixed valve setting. The system watches a live signal that reflects real load, such as:

  • how long requests wait in line
  • how fast work arrives versus how fast it clears
  • the slowest response times
  • the error rate

It compares that signal to a target and gently adjusts how much to admit. It steers by the trend, not just the latest reading, so it makes small corrections instead of overreacting: a dimmer switch, not a hammer. And because it reacts to what is actually happening, it keeps up with changing capacity on its own, with no one editing a config.

THE LOOP, NOT A NUMBER
A four-step loop: measure a signal, compare to target, adjust admission, and measure again.
A closed loop replaces the fixed number: it watches a live signal, compares it to a target, adjusts how much to admit, and then measures the result and does it again.

Two ideas sit underneath it:

  • measure the system you have, not the one you planned - the loop reacts to real, current behavior, not to a capacity someone estimated up front
  • when several controllers share a resource, use one loop, not many - many loops on the same thing fight and swing back and forth; one loop that reads all the signals makes one steady decision

One limit is built in: the loop only reacts after it sees a change, so a sudden spike is absorbed by its response time. That is why it works alongside client-side backoff and jitter, which slow the incoming flood, rather than replacing them. And the loop decides how much to admit; deciding which requests to drop first, when it has to shed, is a separate job for priority-aware load shedding.

When it applies

01The service's real capacity keeps changing. When capacity shifts with the traffic mix, the hardware, or a downstream's health, any fixed limit needs constant retuning to stay right.
02You want to fail gradually, not all at once. A control loop replaces a sudden, all-at-once shutoff - and the retry storm that follows - with a smooth, gradual pullback as load rises.
03You want limits set from real behavior, not guesses. Let the loop derive how many requests to run at once, how long to queue, or how fast to accept from live latency and errors, instead of a number picked up front.
04Several overload signals need to act as one. If signals like local pressure, replication lag, and tenant skew each drive their own limiter, those limiters fight each other. One loop that reads them all makes a single, clear decision instead.

Tradeoffs

The tuning does not disappear, it moves. You stop hand-setting a limit per service, but now the loop itself has to be made stable - not too jumpy, not too slow - and that calibration is done once at the platform level, usually over several tries.
A loop is harder to reason about than a number. 'Why was this request dropped?' has a one-line answer under a fixed limit, but only a controller-state answer under a loop. Operators need a clear view into what the loop saw and decided, or they will not trust it during an incident.
Different signals have to be put on the same scale. A loop that blends latency, lag, byte counts, and request counts needs them in comparable units, or one badly scaled signal will dominate and throw off the decision.
The loop always reacts a step late. Because it measures results, it can only respond after load has already changed, so a sudden spike is ridden out by its reaction time - which is why it pairs with client-side backoff and jitter rather than replacing them.

The same move, 5 ways

Every row is a production system that bet on this pattern — the note says how, in that system's own terms.

Uber
Uber Engineering
2026
Instead of a fixed threshold that snaps fully open or fully shut, a controller adjusts how much to admit smoothly and continuously, using live latency and error signals as feedback (the same idea as a thermostat holding a temperature). Cinnamon's PID controller is the worked example, and its BYOS design extends the same loop to any overload signal (commit lag, write bytes, memory pressure), so what used to be several competing limiters becomes one coherent decision. Read the breakdown →
Netflix
Netflix Technology Blog
2024
Netflix's adaptive concurrency limits are a feedback loop straight out of TCP congestion control: the limit is recalculated each sample from how much responses have slowed (newLimit = currentLimit x gradient + queueSize, where the gradient is best-case latency divided by current latency), with no manual tuning and no central coordinator. When things are fast the limit grows; when they slow it shrinks. The prioritized shedding sits on top and decides which requests to drop once that self-discovered limit is hit. Uber's Cinnamon runs a similar loop with a PID controller over queue and latency signals; both are the same idea of measuring the system you actually have instead of hard-coding a fixed threshold. Read the breakdown →
DoorDash
DoorDash Engineering Blog
2023
Third company. The adaptive concurrency limit is a feedback loop in miniature (latency rises → limit tightens), and Aperture promotes the same loop to a platform: arbitrary normalized signals — Prometheus metrics, SLO deviation — feed one controller producing coordinated actuation, the architecture Uber's Cinnamon called BYOS. DoorDash names the anti-pattern this unification prevents: independent local mechanisms whose uncoordinated actions interact badly during exactly the failures they exist to stop. The post's argument is precisely that the local scale of this loop is insufficient. Read the breakdown →
LinkedIn
LinkedIn Engineering
2023
Here the pattern takes a test-and-back-off form: a starting cap based on the load just before the overload, a survivable level held on detector feedback, occasional upward tests, and a doubling wait when a test brings the overload back. The cap is worked out fresh at each overload, because neither the traffic mix nor the kind of overload repeats. No control-theory machinery, just the same closed loop. Read the breakdown →
Stripe
Stripe Engineering
2017
The worker shedder is a feedback loop: it watches how busy the workers are and sheds more when they're overloaded, less when they recover. Its lesson is a tuning scar Stripe published: if it sheds and restores too fast, it oscillates, dropping traffic, seeing health return, restoring it, and immediately overloading again ('I brought it back! Everything is awful!'). The fix is to move slowly in both directions, which trades some reaction speed for stability. Uber and DoorDash build the same kind of loop with more formal control math (PID, AIMD); Stripe found the right damping by trial and error and, usefully, wrote up the failure it hit on the way. Read the breakdown →

Often used together

Patterns sharing breakdowns with this one — derived from co-occurrence, threshold ≥2 shared.

Problems this pattern answers

The walls where its breakdowns live — each opens the cross-company comparison.