Solve It Once: Cadence and the Case for the Central Workflow Engine

Cadence is Uber's answer to a question every large company faces: where should the durable execution of async workflows live? Uber built one central workflow platform and made it serve over a thousand services, from the most critical (which Uber calls T0) down to the least (T5). It runs 12 billion executions and 270 billion actions a month at 99.9% availability, with workflows written as ordinary code in normal programming languages rather than in a special configuration format. The bet is that the engine is built once while every workflow is written fresh, so the hard parts belong in the central workflow platform, not in each team's code. Uber's 2021 internal survey found teams writing 40% less code for the same functionality. The costs are the platform's to carry. A team of 20 runs a shared dependency the whole company leans on. Keeping one heavy team from slowing down the others is constant work. And the shared database underneath caps how much the whole platform can handle.

Interactive

Crash a worker, take down the orchestrator, ship an engine bug, unleash a noisy service, against the embedded and the central architecture side by side, and read who pays for each.

Open the visualization ↓

Problem

Strip a modern distributed service down and very little of it is the actual feature. Around a small core of unique business logic sits a large ring of machinery that almost every service needs and that is not specific to any of them:

  • saving state so it survives a crash,
  • queues, timers, and retries,
  • failure recovery and failover,
  • scaling and traffic management,
  • observability.

Only that small core, plus its tests, is truly specific to the service; everything else is common. At Uber's scale the commonality is the problem. Thousands of services each face the same crash-between-steps ambiguity, and each team builds its own slightly different version of that safety machinery out of databases, queues, and scheduled jobs. The duplication costs engineering time, and worse, it produces a thousand independently maintained versions of the hardest guarantees in distributed systems.

WHAT ACTUALLY DOESN'T SCALE
A small business-logic core ringed by common machinery, rebuilt across a thousand services, so the duplicated ring does not scale.
Only a small core of a service (its business logic and tests) is truly its own. Around it sits a ring almost every service needs: state, queues, timers, retries, failover. Across a thousand services, that duplicated ring does not scale.

The usual way to consolidate this is a workflow engine driven by a special configuration format that declares the order of tasks and their dependencies. That trades one problem for another. Such a format keeps the engine itself simple, but Uber argues it puts the simplicity on the wrong side: the engine is built once, while a unique workflow has to be written for every use case. As the use cases grow, the format either limits the workflow or grows so complicated it stops being practical. Engineers think in programs, not in configuration. And given the range of things that are really workflows (sign-up flows, batch jobs, scheduled jobs, model training, coordinating microservices), Uber states it plainly: any program is a workflow.

So the question is not whether to consolidate durable execution. It is what shape the consolidation takes, and who carries its risks. A central engine is a shared dependency: every service that leans on it inherits its uptime, its capacity limits, and its neighbors. That shared fate is exactly why some teams instead embed the engine as a library inside each service, keeping one service's failure to that service. Uber took the other branch and spent six-plus years making the central bet hold.

Solution

A Cadence workflow is an ordinary program in a normal language (Go, Java, and others through client libraries), written against Cadence's APIs with few additional constraints. The platform makes that program durable and fault-tolerant without the author having to think about it:

  • state survives crashes,
  • steps retry automatically on failure,
  • timers fire across days or months,
  • and the workflow resumes from where it left off.

The Cadence developer surface covers the whole lifecycle: start a workflow, signal it, schedule it (including cron-style repeats), terminate or cancel it. Workflows can nest as parent and child to compose services, searchable tags let you find one among billions, and a UI shows each workflow's inputs, outputs, call timeline, and relationships for debugging.

BUILD THE ENGINE ONCE
One central Cadence platform of about 20 engineers serves 1,000+ services T0-T5, so teams write 40% less code.
Cadence's bet: build the hard machinery once, in one platform run by about 20 engineers, and let each team write only its own workflow as code. Over a thousand services run on it, and Uber found 40% less code.

The platform side is where the central bet is paid for. Being a shared dependency that the most critical services rely on, Cadence has to ship a lot of machinery:

  • built-in rate limits, hot-shard detection, and per-team usage tracking, to keep one heavy team from hurting the rest,
  • load balancing and scaling at every level (the cluster, the team, the individual workflow, and the queues underneath),
  • configurable behavior on failure: automatic retries, failing over to another region, and resets that roll a workflow back to a healthy point after a bad deploy.

Two more pieces exist because the workflows keep running while the code changes. Versioning lets new code run safely against workflows already in flight, and a testing setup that replays and shadows real production workflows catches mismatches before they ship.

The operating model is deliberate: major features are built and run inside Uber for months, then released as open source, and every release of Cadence is backward compatible. A core team of 20 engineers carries the platform. Other Uber teams contribute to Cadence and to the systems it runs on (Cassandra, Elasticsearch, Kafka, MySQL), and outside companies that run it (DoorDash, HashiCorp, Coinbase) maintain their own Cadence teams.

The results are the argument. Over a thousand services at Uber, from the most critical downward, run on Cadence: long-running workflows, synchronous request handling, coordinating microservices, batch jobs, scheduled jobs, single-instance jobs, data pipelines, and model training. The platform runs 12 billion workflow executions and 270 billion actions a month at Uber alone, has grown 100% year over year for several years, and guarantees 99.9% availability. Uber's 2021 internal survey found teams writing 40% less code to build the same functionality, the number that makes the whole consolidation defensible.

SAME PROBLEM, TWO TRADE-OFFS
Embedded: an engine per service, no shared failure, but fixes roll out everywhere. Central: one engine, one team, shared ceiling.
Two answers to one problem. Embedded: the engine is copied into every service, so no shared failure, but a fix rolls out everywhere. Central: one engine and one team, so a fix ships once, but all services inherit one ceiling.
12 billion
workflow executions per month at Uber alone
40%
less code to implement the same functionality (2021 internal survey)

Tradeoffs

  • The central engine concentrates the very risk that the embedded approach exists to avoid. A thousand services, from the most critical down, share one platform's fate. Its 99.9% availability is now a ceiling that everything above it inherits, and a bad day for Cadence is a bad day for rides, payments, and every pipeline in between. Uber's answer is not to deny the coupling but to invest heavily in managing it. That means a dedicated reliability year in 2022, failing over between regions, careful capacity management, and a platform team whose whole job is keeping the shared dependency boring. The real choice is not who understood the risk. It is whether an organization would rather operate one hard thing extremely well, or spread the hardness into every service.
  • Sharing one platform among many teams means guarding against one team hurting the others has to be constant work, not just an occasional incident. When a thousand teams share the same queues, shards, and database, one team's burst (a million timers firing at once) slows every other team down. How much of the platform exists for this shows in its feature list: built-in rate limits, hot-shard detection, per-team tracking, and capacity managed by rate limits. The isolation that an embedded, one-engine-per-service design gets for free, a central engine has to build, tune, and operate forever.
  • Writing workflows as ordinary code comes with one strict rule: the code must be deterministic, meaning it makes the exact same decisions every time it runs. Because durability comes from re-running the workflow and expecting those same decisions, anything that could vary between runs (like reading the clock or a random number) can now cause a wrong result, not just messy code. That includes even routine code changes made while workflows are still in flight. The platform grows machinery to match: versioning for behavior changes against running workflows, and replaying and shadowing real traffic to catch mismatches. There is even a roadmap item to catch these problems while the code is being written, which quietly admits that today they are only caught later, once the code is running. The 'any program is a workflow' pitch is real, but the program lives under constraints ordinary programs do not.
  • Backward compatibility slowly turns into a burden. Every Cadence release has stayed backward compatible for six-plus years, which is exactly what a thousand dependent services require. It also means mistakes in the API cannot be fixed, only worked around, which Uber concedes is a challenge when problems surface in new APIs. The planned V2 branch exists precisely because some changes cannot be made compatibly. A central platform does not just carry every service's load; it is also held back by how hard all of them are to change.
  • The database under the platform is the limit on everything. Uber says it plainly: database load 'seems to be the bottleneck most of the time,' capacity is managed by rate limits, and the efficiency roadmap (spreading load evenly by design, new workflow modes, new capacity controls) is a plan to win back some capacity. Centralizing a thousand services' durable state pours all their write load into one storage system. The same consolidation that removes duplicated machinery also concentrates the scaling problem, which the platform must now solve on everyone's behalf.
  • A 20-engineer platform team is the visible price of 'solve it once.' The 40%-less-code figure is measured across the teams that use Cadence. The full cost also includes that permanent, growing core team, other Uber engineers who maintain the systems Cadence runs on, and, at partner companies, dedicated Cadence teams of their own. The economics favor the platform at Uber's scale and variety of use cases, which is the honest boundary of the lesson. The central bet is not that platforms are free; it is that at enough scale, the duplicated alternative costs more.
99.9%
availability guarantee: the ceiling every service inherits

Patterns in this article

  • Durable Workflows

    Cadence delivers the classic durable-workflows guarantee: multi-step programs that survive crashes and resume with completed work intact, built on persisted state, replayable execution, and versioning for changes made while workflows are in flight. The same guarantee shows up whether the engine is central like Cadence or embedded in each service. That recurrence across opposite designs is the strongest evidence the pattern is real: the guarantee is the constant, and where the engine lives is the variable.

  • Embedded vs Centralized Orchestration

    This article is the central pole of the pattern; an embedded engine like Airbnb's Skipper is the other. The embedded design puts the engine in each service as a library: no shared dependency, and a failure contained to one service, at the cost of shipping and operating the engine everywhere. The central design, Cadence, is one platform with one team and one availability promise shared by a thousand services. The cost is policing noisy services, being held back by backward compatibility, and concentrating all the load onto one shared database. Same problem, opposite answers, each right for its constraints.

Also solving this

Other systems in behindscale's Partial completion under crashes class: