Pattern · seen in 3 breakdowns across 3 companies

Logical–Physical Migration Split

Do a risky migration in two steps: first make the system act as if the data has already moved (mistakes revert in seconds), then actually move it, the step you can't easily revert.

The mechanism

The pattern at its core: the same bug turns up at the same moment either way, but reverting it costs seconds or hours depending on whether the data has physically moved yet.

LIVE ARTIFACTTHE PATTERN, GENERALOPEN FULL SCREEN ↗
THE IDEAMost migration risk is in the new behavior, not in moving the data. If you move the data first and a bug shows up under traffic, reverting means migrating it all back - slow and costly. If you first make the system act as if the data has moved, without touching it, the same bug reverts with a config change in seconds. So you meet the risk where reverting is cheap, then move the data once the new behavior is proven.
WHAT TO TRYRamp traffic onto the new arrangement both ways. The same query bug shows up at 40% either way; migrating physically first makes reverting a slow data move, while doing the logical step first makes it an instant config change, at a fraction of the user impact.

Ramp onto a new arrangement and hit the same bug either way - then watch reverting cost seconds or hours.

Definition

A migration here means moving data from an old arrangement to a new one - say, from one big table to many smaller shards. Doing it in one go is risky, so split it into two steps, one easy to revert and one hard:

  • logical - first, make the whole system behave as if the data has already moved, using database views, routing rules, proxies, or feature flags, while the data itself stays exactly where it is; to every client, the migration already looks done
  • physical - only after that logical step has proven itself, actually move the data; this is the step that is hard to revert, and it now runs against an arrangement that is no longer new or untested

The whole point is how cheaply you can revert. In the logical step nothing has physically moved, so reverting is just a config change that takes seconds, and you can ramp it up slowly under real production traffic. The physical step is the one you can't easily revert.

TWO STEPS, NO RETURN
Two steps separated by a point of no return: a reversible logical step, then a hard-to-revert physical step.
In the logical step the system acts as if the data has moved while it stays put, so reverting is a quick config change. Only once it is proven do you take the physical step, where the data moves.

The insight behind the split is that most of the risk is in the new behavior, not in moving the data. Wrong routing, queries the new arrangement can't handle, or app code that assumes a certain transaction or ordering behavior - these cause incidents, and every one of them can be tested without moving a single byte. If you rehearse them while reverting is still cheap, the surprises you couldn't have predicted become bugs you have already found and fixed. The physical step is then left with only the risks that genuinely need the data to move before they appear. You can push this further with shadow traffic: send a copy of live requests through the new arrangement and compare its answers against the current one, so production becomes your test without depending on the result.

The same idea shows up anywhere you can rehearse cheaply before a step you can't reverse:

  • expand and contract - add the new column or table first, before any code depends on it, and remove the old one only later
  • branch by abstraction - route calls through an interface so you can quietly swap what is behind it, then switch once the new one is proven
  • dark reads and shadow traffic - run the new system's reads alongside the old one and compare the answers before you switch over

The common discipline is order: do every step you can revert before the one you can't, never after.

When it applies

01The physical move is expensive or impossible to revert. Storage migrations, shard splits, and database swaps, where once the data has moved there is no easy going back.
02You can express the new behavior over the old data. Routing, which queries work, transaction rules - all of it can be put behind a layer on top, without moving anything yet.
03Real production traffic is the only honest test. Nothing else exercises the new behavior the way live traffic does, so you need a safe way to try it there.
04You need to ramp up slowly and revert instantly. When a gradual ramp-up and a near-instant revert are hard requirements, not nice-to-haves, this split is how you get both.

Tradeoffs

The extra layer has a cost of its own. Views, proxies, and flags add overhead and can change the very behavior you are trying to test, so you have to measure it, not assume it is free.
Two worlds run at once during the switch: more configuration to manage, more that can break, and a longer overall timeline.
The cheap revert ends the moment the physical step runs. The split changes when you take the risk, not whether you take it - the hard part is still there, just later.
You only prove what your traffic actually exercises. Any behavior the rehearsal didn't hit is still unknown when the physical step runs.

The same move, 3 ways

Every row is a production system that bet on this pattern — the note says how, in that system's own terms.

Figma
Figma Blog
2024
Postgres views made the sharded topology real to every client while the data stayed on one host; feature flags ramped traffic and rollback was a seconds-fast configuration change. The one-way physical failover ran only after the sharded world had already proven itself under live production traffic. Read the breakdown →
GitHub
The GitHub Blog
2021
GitHub's entire plan is the split's first stage made rigorous: schema domains and SQL linters make the application behave as if partitioned (with violations failing CI) while every byte still sits on mysql1. Only a domain with zero violations earns a physical move. Where Figma gated the physical stage on a feature-flag ramp, GitHub gates it on a linter reaching zero: the same pattern, enforced by the build instead of by traffic. Read the breakdown →
Airbnb
Airbnb Engineering
2015
This is the same pattern a third company reached from a different direction. Airbnb's phase one proves the separation is true in the code before any data moves. Every cross-table join is removed or brought into the application, database permissions are revoked to turn discovery into enforcement, and pipelines are repointed, all before the irreversible promotion. Figma rehearsed the new behavior behind views and flags; GitHub enforced virtual schema domains for years; Airbnb compressed the same idea into a two-week preflight. Read the breakdown →

Problems this pattern answers

The walls where its breakdowns live — each opens the cross-company comparison.