Pattern · seen in 3 breakdowns across 3 companies
Logical–Physical Migration Split
Do a risky migration in two steps: first make the system act as if the data has already moved (mistakes revert in seconds), then actually move it, the step you can't easily revert.
The mechanism
The pattern at its core: the same bug turns up at the same moment either way, but reverting it costs seconds or hours depending on whether the data has physically moved yet.
Ramp onto a new arrangement and hit the same bug either way - then watch reverting cost seconds or hours.
Definition
A migration here means moving data from an old arrangement to a new one - say, from one big table to many smaller shards. Doing it in one go is risky, so split it into two steps, one easy to revert and one hard:
- logical - first, make the whole system behave as if the data has already moved, using database views, routing rules, proxies, or feature flags, while the data itself stays exactly where it is; to every client, the migration already looks done
- physical - only after that logical step has proven itself, actually move the data; this is the step that is hard to revert, and it now runs against an arrangement that is no longer new or untested
The whole point is how cheaply you can revert. In the logical step nothing has physically moved, so reverting is just a config change that takes seconds, and you can ramp it up slowly under real production traffic. The physical step is the one you can't easily revert.
The insight behind the split is that most of the risk is in the new behavior, not in moving the data. Wrong routing, queries the new arrangement can't handle, or app code that assumes a certain transaction or ordering behavior - these cause incidents, and every one of them can be tested without moving a single byte. If you rehearse them while reverting is still cheap, the surprises you couldn't have predicted become bugs you have already found and fixed. The physical step is then left with only the risks that genuinely need the data to move before they appear. You can push this further with shadow traffic: send a copy of live requests through the new arrangement and compare its answers against the current one, so production becomes your test without depending on the result.
The same idea shows up anywhere you can rehearse cheaply before a step you can't reverse:
- expand and contract - add the new column or table first, before any code depends on it, and remove the old one only later
- branch by abstraction - route calls through an interface so you can quietly swap what is behind it, then switch once the new one is proven
- dark reads and shadow traffic - run the new system's reads alongside the old one and compare the answers before you switch over
The common discipline is order: do every step you can revert before the one you can't, never after.
When it applies
Tradeoffs
The same move, 3 ways
Every row is a production system that bet on this pattern — the note says how, in that system's own terms.
Problems this pattern answers
The walls where its breakdowns live — each opens the cross-company comparison.