Pattern · seen in 1 breakdown across 1 company

Dead Man's Switch

Detect failure by watching the system's steady 'I am alive' heartbeat, and raising the alert when that heartbeat stops - because a system that has died silently cannot report anything.

The mechanism

At its core: a healthy system sends a steady heartbeat, and an outside observer just counts it. When the beats stop, that silence is the alert - which catches the failures a system can never report itself, because it has already died or been cut off.

LIVE ARTIFACTSILENCE IS THE ALERTOPEN FULL SCREEN ↗
THE IDEAMost monitoring waits for a failure signal - an error, a failed check - and raises an alert when it sees one. But a system that has crashed, hung, or been cut off from the network cannot send anything, so the loudest failures make no sound at all. A dead man's switch inverts this: the healthy system emits a steady heartbeat, an outside observer counts the beats, and when they stop, the silence itself is the alert. This also settles the endless 'who monitors the monitor' question - the final observer only has to count messages, so it can be trivially simple and live somewhere the outage cannot reach.
WHAT TO TRYWatch the heartbeat, then kill the system so it dies silently. The 'wait for an error' detector never fires - the dead system sent nothing - but the dead man's switch sees the heartbeat drop and alerts.

A crashed system can't send an error - but a dead man's switch hears the heartbeat go silent and raises the alert anyway.

Definition

A dead man's switch spots failure by watching for a signal to stop, not for an error to appear. A healthy system sends a steady heartbeat - a regular message, a timestamp it keeps refreshing, a probe that keeps passing - and an outside observer just counts them. When the heartbeat stops, the observer raises the alert. The silence itself is the signal.

SILENCE IS THE SIGNAL
Left, waiting for a failure signal misses a silent death; right, watching for a missing heartbeat catches it.
Waiting for a failure signal misses a silent death - a crashed or cut-off system sends nothing; a dead man's switch watches for the heartbeat to stop, so the silence itself raises the alert.

A monitor is itself a system that can fail, which leads to a problem with no end. If a monitoring service watches your app, what watches the monitoring service? Adding another monitor on top just moves the question up a level, forever - and the one at the very top can still die quietly. A dead man's switch escapes this by flipping the question. Instead of asking 'did something break?', which needs a working detector to answer, it asks 'is the system still saying it is alive?', which needs nothing more than the ability to count messages.

The final observer is kept as dumb as possible - a simple alert that fires on a low message rate, a scheduled job checking a file's timestamp, a load balancer expecting a certain reply. And it has to live in a different failure domain from the system it watches; a watcher that shares infrastructure with what it monitors gives no protection, because the same failure takes them both. The whole pattern rests on the observer surviving when everything else has died.

The same idea shows up well outside software. In each case, something has to keep actively proving it is okay, and the moment it stops, that is the warning:

  • a device that reboots itself if its software stops checking in
  • a train that brakes on its own if the driver stops pressing a pedal
  • a machine that runs only while two operators each hold a button down
  • a medical alert that calls for help if someone misses their daily check-in

When it applies

01Your monitoring needs watching too, without an endless chain. If you add a monitor to watch your monitor, then another to watch that one, it never ends - a dead man's switch replaces the whole chain with one check for silence.
02The failure is silent - nothing gets reported. Scheduled jobs that never run, processes that vanish without an error, a network split where the broken side cannot call for help: none of these send a failure signal, but all of them stop the heartbeat.
03You need to know a safety system actually runs, not just exists. For alerting pipelines, backups, or replication, a heartbeat proves the thing is really running - configuration alone does not.
04The failure can take out the system's own monitoring. When a whole region can go down, a process can stop dead, or the network can split, the system's built-in telemetry dies with it - so you need an outside heartbeat to notice.
05The failure could take the detector down too. Put the heartbeat sender and the watcher on independent infrastructure, so whatever kills the system cannot also silence the thing meant to notice.

Tradeoffs

It only tells you alive or dead, nothing in between. The heartbeat is present or absent, so it says nothing about how degraded the system was on the way down - that needs separate, finer monitoring.
How often it beats trades speed against noise: a fast heartbeat catches failures sooner but adds traffic and can fire by mistake when the network hiccups for a moment; a slow one is cheaper but lets an outage run longer before it pages.
A missed heartbeat is ambiguous. It could mean the system is down, or just that the network path between it and the observer is down - the observer often cannot tell the two apart, which makes for noisy alerts during network trouble.
The observer has to sit somewhere the failure cannot reach. A watcher sharing infrastructure with the system it watches protects nothing, since the same failure takes both - the whole pattern depends on the observer outliving everything else.

The same move, 1 ways

Every row is a production system that bet on this pattern — the note says how, in that system's own terms.

Airbnb
Airbnb Engineering
2026
The heartbeat chain (a monitor that emits a steady signal, an external service that counts the signals, and an alarm that fires when the count drops) is the textbook form of the pattern. It watches for the absence of a good signal rather than the presence of a bad one, which turns the endless 'who monitors the monitor' problem into one simple check for silence. Read the breakdown →

Problems this pattern answers

The walls where its breakdowns live — each opens the cross-company comparison.