The Drain Button: Slack's Migration to a Cellular Architecture
On June 30, 2021, a flaky network link in one AWS availability zone degraded Slack for everyone. The incident review's real question was why a failure in just one zone reached users at all. The answer was a gray failure: different parts of the system disagreed about what was actually up. A single user request fans out into hundreds of internal calls that all have to succeed. And Slack's main database keeps each piece of data on one machine that must be reachable to save changes, so when the flaky link cut those machines off, those saves failed. A partial, confusing failure in one zone turned into errors everywhere. Rather than try to detect these gray failures automatically, Slack spent 1.5 years moving its most important user-facing services to cells. Every service runs in every zone but talks only to services in the same zone, turning each zone into a self-contained cell. The result is a single control that pulls all traffic out of a troubled zone in seconds, 1% at a time, without needing anything inside that zone to still work.
Inject the June 30 gray failure into both layouts and watch who sees errors, then hit the drain button and step a zone's weight down to zero.
Problem
Slack serves users from around the world at its edge, but most of its real computing lives in a handful of availability zones inside one AWS region, us-east-1. An availability zone is a separate datacenter, and the cloud provider designs them so that a problem in one is unlikely to hit the others at the same time. The whole point is that a service spread across several zones should be more reliable than any single zone. The June 30, 2021 incident broke that promise. At 11:45am PDT, a network link joining one zone to several others started faulting on and off, slowing and breaking connections between Slack's servers and hurting service for customers. The cloud provider took the faulty link out of service, put it back after automated checks passed, and then removed it permanently when it failed again that evening.
The incident review asked the uncomfortable question: why did this reach users at all? Several things pile up:
- A single request from a user (say, loading the messages in one channel) can fan out into hundreds of internal calls, and every one has to succeed for the user to get a correct answer.
- Slack's systems watch for failed backends and route around them, but a server has to fail a few times before it can be marked bad and skipped.
- Slack's main database, Vitess, keeps each slice of data on one machine that must be reachable to accept a change. If the network cannot reach that machine, changes to that slice fail until it comes back or a backup takes over.
The June 30 outage was a gray failure: different parts of the system held different views of what was healthy. Inside the hurt zone, local machines looked fine and far-away ones looked down; from outside, the hurt zone looked down; and even two clients in the same zone disagreed, depending on whether their traffic happened to cross the broken hardware. That is a lot of confusion to ask a distributed system to sort out while it is trying to do its real job: serving messages and cat GIFs. To the engineers on the incident, though, the cause was obvious: nearly every graph, grouped by which zone was involved, pointed at the same one. They realized that if they had one way to tell every system 'this zone is bad, send traffic elsewhere,' they would have used it instantly. So Slack decided to build exactly that.
Solution
Slack's fix is a control they call the drain button. Draining a zone means pulling all user traffic out of it; undraining means letting traffic back in. The design goals for it are demanding:
- Fast: pull as much traffic as possible out of a zone within five minutes. Slack's reliability promise allows under an hour of downtime a year, so mitigation has to act quickly.
- Harmless: draining must not itself cause errors, because it is meant to be tried before anyone knows the cause. If draining caused errors of its own, an operator could not safely try it during an incident and reverse it if it did not help.
- Incremental: draining and undraining happen in small steps, down to sending just 1% of traffic back into a zone to check whether it has really recovered.
- Self-sufficient: the mechanism must not depend on anything inside the zone being drained, so it still works when that zone is completely offline.
The obvious approach, teaching every service to drain itself, falls apart because Slack's systems are so varied. Three problems in particular:
- The user-facing services are written in four languages (Hack, Go, Java, and C++), so it would take four separate implementations.
- The ways services find each other are also mixed (Envoy, Consul, and plain DNS), and plain DNS has no notion of a zone or a partial drain at all.
- Heavily used open-source systems like Vitess would force an unpleasant choice: keep a private copy of the changed code, or do the extra work to get the changes merged back into the shared open-source project.
The strategy Slack chose is called siloing. A siloed service only takes traffic from within its own zone and only sends traffic to servers in its own zone. In effect, each service becomes several separate copies, one walled off inside each zone, and the zone itself becomes the cell. That leads to the key simplification: to pull traffic out of every siloed service in a zone, you only have to stop user requests from entering that zone. With no new work coming in, the services inside wind down on their own. A failure in one zone stays in that zone, and traffic routes around it based on a single decision made at the front door.
Siloing pushes the whole traffic-shifting problem into one place: the systems that route user requests into the core services in us-east-1. Slack had already spent years moving its edge load balancers onto Envoy, a widely used proxy, controlled by an in-house system called Rotor. The drain button then falls out of two standard Envoy features that let you split traffic across targets by adjustable weights. Draining a zone just means telling the edge load balancers, through Rotor, to reweight that zone's share. Set the weight to zero and Envoy finishes the requests already in flight and sends all new ones to the other zones. The change spreads through the system in seconds, no requests are dropped, and the weights allow 1% steps. Crucially, the edge load balancers live in entirely different regions, and their control system is copied across regions, so nothing about draining depends on the sick zone still working. Slack's own graph of traffic per zone shows the payoff: sharp, near-vertical drops as traffic steps out of one zone into the other two. Getting here, siloing Slack's most important user-facing services, took 1.5 years, and Slack said the details of siloing internal services, the services that cannot be siloed, and the operational changes would come in later posts.
Tradeoffs
- The button works by giving up on full automation. Gray failure defeats the automatic failure detection Slack already had: the components cannot agree a failure exists, so they cannot converge on it. Handing the decision to a human is a deliberate step back from self-healing, made because a responder can see what the machines cannot, by pooling evidence from many views. The cost is that mitigation now waits on a human, with human delay and the burnout of being paged again and again. The design goals (five minutes, 1% steps, safe to try) exist precisely to make that human action cheap enough to take on a hunch rather than proof.
- Siloing trades shared capacity for contained failure. When each service only talks within its own zone, the healthy zones can no longer lend spare capacity to absorb a spike in a struggling zone mid-request. Each zone must now be sized to carry its own share plus the extra traffic drained from a failed neighbor. The old design got capacity-sharing for free and paid for it in blast radius; the cellular one flips that trade. Keeping that spare capacity in every zone is a permanent, quiet cost across the whole system.
- The drain works no matter what the problem is, which is its strength, but that also makes it blunt. A drain moves an entire zone's traffic whether the failure touched one service or all of them. That is the article's main example of a generic mitigation: something you can use before you know the root cause. The bluntness is the point, because during a gray failure you do not know the root cause. But every drain over-does it, moving healthy work along with the sick, and the team has to be willing to pay that cost every time something looks wrong.
- Doing all the draining at the front door puts a lot of trust in that front door and the system that controls it. The whole design sidesteps the per-language, per-discovery complexity by making one mechanism, reweighting at the edge, do the job. That makes the edge load balancers and Rotor a single point of failure, the one part that must never get it wrong. The design accounts for this: the edge runs in other regions and its control system is copied across regions, so the drain path is built to be more available than the thing it drains. It is a smaller, better-defended single point of decision, not the absence of one.
- Not everything can be siloed, and consistency is the reason. Strongly consistent databases like Vitess keep each slice of data on one machine that must be reached to save a change, and that machine lives in some particular zone. A service that must reach that one machine cannot treat the zone boundary as a solid wall, which is why Slack immediately promised a follow-up on the services that cannot be siloed. The cell boundary works for ordinary requests that any machine can serve, but not for writes that must reach one specific machine, the stateful data that lives in a single place. That is the real limit of the design, and the place where the next incident will look different rather than vanish.
- The migration cost 1.5 years, and what it bought is undramatic: a way to mitigate, not a way to prevent. Nothing about cells stops the next flaky network link; the work buys the ability to stop caring about it quickly. At a target of under an hour of downtime a year, that is a defensible bet: fast, general mitigation is worth more than trying to automate away every root cause. But it is still a bet, that reliability effort is better spent on recovering fast than on preventing failures, and placing it meant restructuring most of Slack's critical services.
Patterns in this article
- Cell Architecture
Slack's cells are availability zones: every service runs in every zone, and no service talks across zone boundaries, so each zone is a unit you can drain on its own. It is a clean illustration of the general idea: a cell is whatever unit you can afford to lose, sized so that losing it is a routine event rather than a crisis.
- Fault Isolation
Siloing changes a zone from a piece of every request's fan-out into a contained failure zone: a failure inside one zone shows up only in that zone's frontends, instead of in all of them. The June 30 incident is the negative proof: before cells, one link's gray failure was everyone's outage.
- Generic Mitigation
The drain button is a clear example of a generic mitigation, which the article names directly: a fix you can apply while the root cause is still unknown. It is safe to try because it adds no errors of its own (drain, watch, undrain if it does not help). The four design goals, fast, harmless, incremental, and independent of the failing zone, are basically the requirements for that kind of mitigation, written down.
Also solving this
Other systems in behindscale's Gray failure defeats automatic detection class: