Monitoring Reliably at Scale
Airbnb's monitoring system ran on the same shared infrastructure it was supposed to watch, so when that infrastructure broke, the monitoring broke with it, going blind at the exact moment engineers needed it most. This is a circular dependency: the tool meant to detect an outage depends on the very thing that is failing. Airbnb rebuilt its monitoring to remove that loop, in three layers. First, it moved the monitoring onto its own dedicated machines, still run by the platform team but not shared with the services being watched. Second, it gave monitoring data its own network path, separate from the shared networking layer that every other service uses. Third, it added a 'dead man's switch': the monitoring sends out a steady heartbeat, and if that heartbeat ever stops, the silence itself is what pages the on-call engineer. The one rule running through all of it: never let your safety net depend on the thing it is supposed to catch.
Break each layer, in the old architecture and the new, and see who gets paged.
Problem
When an incident hits, engineers turn to their monitoring to answer two questions: what is broken, and why. At Airbnb, thousands of services run on shared infrastructure, and the monitoring system meant to answer those questions ran on that same shared infrastructure. That created a circular dependency: the monitoring depended on the very systems it was meant to watch. So when the shared infrastructure failed, the monitoring of that infrastructure failed at the same time, and the team went blind at exactly the moment visibility mattered most. The dashboards go dark and the alerts stop firing, and the tools meant to guide the recovery become part of the outage.
The loop is easy to trace, and every step lands back on the same shared foundation. The product services ran on shared Kubernetes clusters (Kubernetes is the system that runs services in containers). The monitoring pipeline ran on those same clusters. The monitoring data traveled over the shared service mesh, the networking layer that connects services. And that service mesh itself ran on the same shared Kubernetes clusters. So a problem anywhere in this shared foundation traveled upward and switched off the alerts that were supposed to warn the team about that exact problem.
There was a second, quieter problem alongside the loop. Monitoring data at Airbnb's scale is far larger than ordinary business traffic, many times over, because every service is constantly pushing monitoring data to a central store, which the shared networking layer was never designed for. Putting both kinds of traffic on that one shared layer meant they competed for the same capacity. A surge of monitoring data could crowd out real user requests and hurt Airbnb.com itself, and a surge of user traffic could crowd out the monitoring, so engineers would lose visibility right when load was highest. The two kinds of traffic had different needs and deserved different priority, but the shared layer treated them exactly the same.
Solution
The fix comes in three layers, each one breaking a different loop.
Compute isolation: give monitoring its own machines. Airbnb faced two extremes, and neither worked:
- Run the monitoring on the shared production clusters. This needs no new setup, but it is exactly the coupling that caused the problem: the monitoring relies on the very thing it watches.
- Run its own separate clusters end to end. This gives full isolation, but it means a small team taking on the deep, constant work of operating clusters, which was not sustainable.
The answer was in between: dedicated Kubernetes clusters just for monitoring, not shared with any product or infrastructure services, but still run and maintained by the platform (Cloud) team. That keeps Kubernetes as a managed foundation while removing the shared point of failure. To keep it safe, the two teams coordinate changes so that only one big change lands at a time, and every change is tried on lower-priority clusters before it reaches the ones carrying real monitoring load.
Network isolation: give monitoring data its own separate network path. Airbnb built a custom entry point for monitoring traffic, based on a proxy called Envoy, that sits entirely outside the shared service mesh. Monitoring data now travels on its own network path, with its own priority, so the shared mesh can fail without taking the monitoring down, and a flood of monitoring data can no longer congest the mesh and hurt user traffic. Running this path on separate machines from the shared compute adds a further margin of safety. Owning this layer also unlocked features the shared mesh could not offer. Airbnb runs over 1,000 services. The custom layer keeps each service separate from the others but manages them all in one place, mapping each service to the right backend, and every request carries a small label that tells the layer which service it is for. That label-based routing spares each service from complex configuration and spreads load evenly. The team can also copy monitoring data to other destinations for testing, and enforce fine-grained access controls, which matters when outside vendors are involved.
Why own the network but not the compute? Because the two needed different things. Kubernetes was already a mature, managed platform, so getting isolation there was just a matter of asking for dedicated clusters, a small addition for the Cloud team. Networking was different: the shared mesh simply could not cleanly separate and prioritize monitoring traffic at Airbnb's scale, and the specific features they needed (strict priority, isolation, custom routing) sat squarely in the monitoring team's own area. Owning that layer gave them the control they needed, and it was far simpler to operate than running Kubernetes themselves.
Meta-monitoring: watch the watchers. Once the core layers were solid, one question remained: how do you know when the monitoring itself is sick? Airbnb runs a separate set of monitoring servers (using a tool called Prometheus) whose only job is to watch the main monitoring stack. To avoid failing together with what they watch, these servers run on machines kept apart from the main stack and spread across separate datacenters. They are also paired so that no monitor and its alert-sender ever sit on the same shared hardware. But that just raises the same question one level up: who watches these watchers? Stacking on yet another monitor would go on forever.
The answer is a dead man's switch: a mechanism built around a steady signal whose absence is the warning. Airbnb keeps an alert rule that fires continuously as long as its monitoring is healthy, sending a steady heartbeat out to an external service on Amazon's cloud. A separate Amazon alarm simply watches how often those heartbeats arrive. If they stop, for any reason at all (the monitor crashed, it stalled, the sender failed, anything else), the alarm trips and pages the on-call engineer. The endless 'who watches the watchers' problem collapses into one dead-simple check for silence. The silence itself becomes the alarm.
Tradeoffs
- Dedicated clusters add coordination work with the platform team: monitoring changes now have to line up with broader platform changes. But the alternative, a small team running its own clusters end to end, was simply too much for its headcount to sustain.
- Owning the network path but not the compute creates a split arrangement: the monitoring team owns the custom traffic layer, the platform team owns Kubernetes. That boundary needs ongoing upkeep as both sides change, but it lets each layer be run by the team best suited to it.
- The dead man's switch is deliberately simple: the external alarm just counts heartbeats. That simplicity is also its limit. It can only catch total silence, not a slow or partial decline. The one thing it tells you is binary, that the heartbeats stopped, so spotting subtler trouble still needs separate monitoring on the stack itself.
- Running a custom traffic layer means the team has to understand and maintain networking infrastructure, a skill set that does not carry over directly from ordinary application work. They take on that learning cost because the alternative, sharing the mesh, leads to worse failures.
Patterns in this article
- Independent Observability
The common thread of the redesign is that no part of a safety mechanism may depend on the thing it protects. Airbnb breaks the loop at three points: dedicated clusters break the compute dependency, a separate network path breaks the data-flow dependency, and an external heartbeat watched from outside Airbnb's own infrastructure breaks the alerting dependency. Each was a hidden loop, invisible in normal operation and dangerous only when it mattered. The lesson is the discipline of mapping every such loop before deliberately cutting it.
- Fault Isolation
Every layer of the redesign is the same fault-isolation idea applied again. Dedicated clusters separate the monitoring pipeline from the product workloads, a separate network path separates monitoring traffic from business traffic, and the dead man's switch separates failure detection from the failing system. One principle, applied at three different layers.
- Dead Man's Switch
The heartbeat chain (a monitor that emits a steady signal, an external service that counts the signals, and an alarm that fires when the count drops) is the textbook form of the pattern. It watches for the absence of a good signal rather than the presence of a bad one, which turns the endless 'who monitors the monitor' problem into one simple check for silence.
Also solving this
Other systems in behindscale's Observer shares fate with observed class: