Skipper: Building Airbnb's Embedded Workflow Engine
Airbnb built Skipper as a workflow engine that ships as a library embedded inside each service, not as a separate central cluster. Durability comes from replay: when a service restarts, the workflow method simply runs again, skipping every action already saved (checkpointed) to the database. A workflow that has to wait uses no compute at all while it waits. If it fails midway, Skipper automatically undoes the steps that already ran, in reverse order. And all of this state lives in the database the service already uses. The whole design is built around one choice: pay almost nothing when things go smoothly, and call on the recovery machinery only when something goes wrong. In production over a year across 15+ use cases, peaking at 10,000 workflows/second.
Step the workflow, crash it anywhere, and watch replay skip the actions that already committed.
Problem
Airbnb runs many multi-step business processes (insurance claims, payments, media processing, infrastructure automation) that span minutes to days. Take a claim: it might validate, run trust-and-safety checks, assess estimates, process a payout, and notify the host. If the server crashes after validation but before the payout, a traditional design leaves the outcome to chance: a timeout that makes the caller retry, running the work twice, or partial state that corrupts what follows. When a process waits for hours or days, being interrupted partway is not a rare edge case; it is the normal thing to plan for.
The industry already has answers, and Airbnb weighed three:
- A dedicated orchestration cluster (systems like Temporal or Cadence): a separate shared cluster that runs the workflows. It is reliable and gives exactly-once guarantees, but it is a new critical dependency, and for Airbnb's most critical services (the ones directly behind user-facing transactions) that was disqualifying: one cluster outage would stop every dependent service from starting or advancing a workflow.
- A cloud-managed workflow service: the same shared-dependency concern, plus vendor lock-in and data-handling constraints.
- A homegrown queue-based system: it avoids the outside dependency, but every team ends up building its own tangle of retries, state, and cleanup.
There is also a second, less obvious problem, and it is common across the industry. When a multi-step process is wired together from queues, scheduled jobs, callback endpoints, and reconciliation scripts, the business logic gets fragmented. There is no single place in the code that says what the process actually does, and every fragment tangles the business rules together with retry timing, duplicate-suppression, and timeout handling. Teams were each rebuilding this, rediscovering the same crash-recovery edge cases, and shipping their own subtle bugs.
Solution
Skipper ships as a library embedded directly inside each service. There is no central cluster. Workflow state lives in the database the service already uses (MySQL or Airbnb's internal Unified Data Store), and the engine runs inside the service on its own threads.
The programming model is built from a few simple pieces:
- A Workflow is a plain Kotlin or Java class whose method spells out the process itself: the order of steps, the conditionals, the waits. It reads like the business process it represents.
- An Action wraps a single side effect (an API call, a database write, a notification). A single `@Execute(checkpoint = true)` annotation saves that action's result to the database after it runs (a checkpoint).
- `@StateField` fields hold the workflow's state.
- `@SignalMethod` lets an outside event (a reviewer's decision, a callback) push data into a running workflow, updating the state fields that `waitUntil` conditions check against.
Durability comes from replay. Skipper runs the workflow method and checkpoints each action's result. When execution resumes (after a signal, a timer, or a crash), Skipper replays the workflow method from the top. Previously completed actions do not run again; they return their saved results instantly, and the method picks up where it left off. One important choice: Skipper saves the state fields directly, rather than rebuilding them by replaying a full history of past events the way some other engines do. There is just the current state and the saved action results. That keeps execution lean, especially for workflows with many signals or long histories, in exchange for giving up some ability to trace back, step by step, exactly what happened.
A workflow that has to wait goes to sleep. When a workflow hits `waitUntil`, its state is written to the database and the thread is freed. The workflow then exists only as a row in the database until a signal arrives or the timeout fires. It uses zero compute and does no polling, whether the wait is seconds or weeks.
When a workflow fails in mid-run, earlier actions have already taken effect. For example, a listing's photos passed their content check, but then the listing never got published, so that check was done for nothing. Skipper makes undoing that work a built-in feature. The `@Compensate` annotation pairs each action with an undo method, and on failure Skipper runs the undos in reverse order (release inventory, refund charges, revert state), walking the system back to a consistent state. You get eventual consistency without database-spanning transactions, written as 'what does undo mean for this action' rather than manual cleanup jobs.
The best part of the design is the happy path. Most engines add overhead on every run: a central orchestrator needs a network round-trip per step to save its result before moving on. Skipper flips this around. At the start it does two things in the database: it creates the workflow instance, and it schedules a delayed timeout task as a safety net. Then the workflow runs entirely inside the service. Actions execute as ordinary method calls on an in-memory queue, checkpoints are batched, and the workflow runs to completion with no further coordination. The delayed task is what makes this safe: if the process crashes mid-run, a scheduler picks the workflow up after a short waiting period (long enough to be sure the original run really died) and replays it; if the workflow finishes normally, the timeout fires harmlessly and is thrown away.
So on the happy path Skipper adds just a few database writes. The engine only steps in when something goes wrong:
- a crash triggers a replay,
- a `waitUntil` hibernates the workflow,
- an error runs the compensations.
Durability is guaranteed, but you only pay for it when you need it, which is what makes an embedded engine practical inside fast, high-traffic services. Production bears this out. Skipper has run for more than a year, powering 15+ use cases across insurance, payments, media, infrastructure, incentives, and wallet teams. At peak it has scaled to 10,000 workflows per second on DynamoDB, which its lean, coordination-free execution is what makes possible.
Tradeoffs
- Workflow methods must be deterministic, meaning they must behave identically every time they run. Given the same inputs, saved results, and state fields, the method has to make the same decisions and call actions in the same order on every replay. So anything that could vary (randomness, clock reads, API calls) must live inside actions, never in the workflow body. Airbnb names this the model's least intuitive requirement for newcomers, and the friction is real: you now have to keep replay in mind whenever you reason about why a run did what it did.
- Actions run at least once, not exactly once. If an action succeeds but the process crashes before its checkpoint is written, the replay runs it again. So each action must be safe to run more than once (idempotent): the guarantee covers the workflow's overall progress, not whether a single side effect happens exactly once, and making sure of that falls to the team using Skipper.
- Changing a workflow over time is the most common friction. Changing a workflow's structure while instances are still in flight can break replay, because the saved sequence of actions no longer matches the new code. Airbnb has versioning patterns (new method versions, traffic migration, deprecation) but explicitly asks for better tooling (automated compatibility checks, migration assistants, runtime versioning) that does not yet exist. Teams that adopt Skipper are stuck with this rough edge for now.
- Seeing what your workflows are doing takes a replay-aware mindset and some custom tooling. Workflow state is just rows in MySQL or UDS, not a ready-made dashboard, so any cross-workflow view has to be built. Worse, debugging a replayed workflow means reading logs whose timestamps and call order reflect the replay rather than the original run. Airbnb flags replay visualization as tooling it wishes it had.
- The embedded model gives up what a central orchestrator provides for free: coordinating across services, and support for many languages. Skipper lives inside one JVM service, so a workflow that spans services becomes API calls and hand-coordinated steps across them, and non-JVM services cannot use it directly. Airbnb is explicit that teams needing either should prefer a dedicated orchestration system. The embedded choice is right precisely when keeping dependencies few matters more than reaching across services.
- Saving only the current state, rather than a full history, keeps Skipper fast and lightweight, and it avoids replaying long histories (the lever behind its throughput). But it gives up the complete, replayable record of everything that happened that event-log-based systems provide for free. Where regulatory or forensic history matters, that is the wrong trade to make.
Patterns in this article
- Durable Workflows
Skipper's Workflow-plus-Action model is a clean embedded version of the pattern: deterministic orchestration logic, side effects checkpointed behind a single annotation, and crash recovery by replaying from the last good state. Shipping it as a library rather than a central cluster changes where durability runs, not the pattern itself, which is why the same pattern shows up in centralized engines too.
- Embedded vs Centralized Orchestration
This is the embedded side of the choice. Airbnb evaluated centralized orchestrators (Temporal, Cadence) and rejected them because a shared cluster is an unacceptable single point of failure for its most critical services. Skipper trades reach across services and languages for failure isolation and zero new infrastructure, and Airbnb is unusually explicit that this is the right trade only when keeping dependencies few is what matters most.
- Hibernation vs Polling
Skipper's waitUntil writes the workflow to the database and frees the thread; the workflow then waits as a row, using zero compute, until a signal or timeout wakes it. This collapses the usual sprawl (a queue consumer, a callback endpoint, a scheduled job, a state table) that the same 'wait hours for approval' logic would otherwise need into a single readable method.
- Atomic Phases
Skipper's checkpointed Actions are this pattern applied inside a single workflow method. Each action commits its result durably, and replay resumes from the last committed point rather than redoing or skipping work. The workflow method is the sequence, and each checkpoint is a phase boundary.
Also solving this
Other systems in behindscale's Partial completion under crashes class: