Strangler-fig is the most-cited pattern in legacy modernisation and the most badly executed. The idea — route traffic incrementally from the old system to a new one, retire the old system piece by piece, no big-bang cutover — survived two decades because it works. What hasn’t survived is the discipline. Most strangler-fig programmes we walk into are not strangler-fig programmes. They are parallel rewrites with a routing layer in front, drifting indefinitely, with neither system in a state where it could be retired.
The reframe that makes them ship: strangler-fig is a cutover strategy, not a refactor. Every architectural decision in the programme is in service of one question — what does it take to switch this slice of traffic permanently? When the team holds that question in mind, the work resolves. When they don’t, it doesn’t.
The “rewrite plus router” failure mode
The shape we keep seeing is the same. A team takes on a legacy modernisation. They stand up a new service. They put a routing layer — an API gateway, an Envoy filter, a Cloudflare worker, a custom proxy — in front of the legacy system. They start moving traffic.
Six months in, the new service handles 30% of read traffic for two endpoints. Writes are still going to the legacy system, because the new data layer’s consistency model isn’t compatible with the legacy assumptions. The team is now maintaining two stacks. Velocity has halved. The legacy system, contrary to the plan, is still receiving feature work — because the business needs the features and the new system can’t accommodate them yet.
This is not strangler-fig. This is two systems, joined by a routing layer, neither of which can be retired. The fig has not strangled anything. The host tree is still alive and growing.
The fig has not strangled anything. The host tree is still alive and growing.
The fix isn’t more engineering investment in the new service. The fix is the discipline that should have been there from the start: every slice of work has a defined cutover, a defined rollback, and a defined retirement of the legacy code path. No work begins on a slice until those three are written down.
Slice on the cutover boundary, not the architectural one
The standard modernisation playbook says to slice by bounded context: identify the domain seams, build the new service around one, route the relevant traffic. Conceptually right. Operationally, it produces slices that are too big to cut over cleanly.
Cutover slicing is finer-grained. The unit isn’t “the customer service” — it’s “the one read endpoint on the customer service that takes a customer ID and returns a profile”. That endpoint can be built, traffic-shifted with a feature flag, observed in shadow mode against the legacy response, and finally cut over with a defined success criterion. Then the next endpoint. Then the next.
The result looks less elegant on an architecture diagram and ships an order of magnitude faster. The new service grows endpoint by endpoint, each one with a known good cutover, each one removing a small piece of legacy that can be deleted.
The unit you measure is “endpoints retired from the legacy system this quarter”, not “endpoints implemented in the new system this quarter”. They are not the same number. A team that conflates them will report progress that doesn’t translate to legacy retirement.
Shadow mode is non-negotiable
The single highest-leverage practice in strangler-fig is shadow mode: every request to the legacy system is duplicated to the new one, the responses are diffed, and the discrepancies are logged. The legacy response is the one returned to the user. The new system runs in the dark.
Shadow mode does three things at once. It tells you the new implementation is correct, against real traffic, without any user risk. It exposes the cases the spec didn’t capture — the long tail of inputs the legacy system handles in ways nobody documented. And it builds confidence in the team and the business that the cutover will be uneventful.
Without shadow mode, the cutover is a leap of faith. The team finds out about edge cases when users complain. The rollback gets used. The next slice is approved more reluctantly. The programme drifts.
// Sketch of a shadow-mode handler at the routing layer.
// The legacy response is authoritative; the new response is observed.
async function handle(req: Request): Promise<Response> {
const legacyPromise = forwardToLegacy(req);
// Fire-and-forget the new service. Do not await before responding.
forwardToNew(req)
.then(async (newRes) => {
const legacyRes = await legacyPromise.clone();
const diff = await compare(legacyRes, newRes, req.path);
if (diff.discrepant) {
await logDiff({
path: req.path,
requestId: req.headers.get("x-req-id"),
fields: diff.fields,
severity: diff.severity,
});
}
})
.catch(err => logShadowError(err));
return legacyPromise;
}
The discipline around the diff log is what matters. Discrepancies are triaged like incidents. A daily diff report is reviewed by the team. A path doesn’t move from shadow to live until the diff rate is below a defined threshold — for read endpoints, typically zero; for write endpoints, defined per-field with explicit business sign-off.
The data layer is the project, not the service
A strangler-fig programme that doesn’t take the data layer seriously is going to fail. The new service can be beautifully implemented, but if it reads and writes the same database the legacy system does, you have not modernised. You have introduced a second client to a system whose schema is still the constraint.
Three honest options exist for the data layer.
The first: keep the legacy database, modernise the access layer only. This is the cheapest path and the right one for some workloads — usually those where the schema is fine and the problem is the application code. The cutover is straightforward; the upside ceiling is lower.
The second: dual-write to both old and new data stores during the migration, with the new store designed properly, and a defined moment when the new store becomes authoritative. This is the most common path and the most operationally complex. The dual-write window is dangerous; reconciliation jobs are mandatory; the cutover of authority is a separate, named event.
The third: change data capture from the legacy database into an event stream, project into the new store, run the new service against the new store, switch writes only at the end. This is the cleanest separation but requires that the legacy system tolerate a CDC connector and that the team is comfortable operating Kafka or equivalent.
Pick one explicitly. Mixing them across slices, or not picking at all, is how you end up with a strangler-fig programme that, two years in, still has the legacy database as the source of truth for everything that matters.
Decommissioning is part of the slice
A slice is not done when the new endpoint is live. It is done when the legacy code path is deleted, the legacy infrastructure is shut down, and the legacy access is revoked.
Most programmes treat decommissioning as a phase at the end. The phase never arrives, because by the time the new system is “complete enough”, the legacy system has accreted three years of cruft from teams that needed to ship features against it. Decommissioning becomes its own multi-month project, scheduled forever for next quarter.
The fix is to make decommissioning the final task of every slice. The story isn’t “build new endpoint”. The story is “build new endpoint, cut over, run for two weeks, delete legacy code, decommission legacy infrastructure”. The slice is not closed until the diff between the legacy repository’s main branch this quarter and last quarter shows the deletion. Code that hasn’t been deleted is still running, even when nobody thinks it is.
This is also where the financial business case for the programme actually closes. The savings come from running fewer systems, not from running a new one alongside the old one indefinitely. A strangler-fig programme that hasn’t decommissioned anything has not yet saved anyone anything; it has only added cost.
Feature freezes are a tool, not a stance
The orthodox strangler-fig advice is to freeze the legacy system. No new features. All work goes to the new system. This works in theory and almost never in practice — because the business has a roadmap, the roadmap doesn’t pause for modernisation, and the team that says “the legacy system is frozen” is the team that gets overruled in the steering committee.
A more workable rule: features go to whichever system is faster to build them in, unless they touch a slice that’s currently mid-cutover, in which case they wait for the cutover to complete or get built in both systems. The priority is keeping cutovers clean, not keeping the legacy system pristine.
This means accepting that the legacy system gets some additional work done to it during the programme. That’s fine. It also means accepting that some slices get built fast and ugly in the new system because the business needed them now. That’s also fine, as long as the cutover discipline is preserved.
The thing you don’t compromise on is the cutover. Everything else is negotiable.
Knowing when to declare the programme finished
A modernisation programme has a quiet failure mode at the end: the team has retired 80% of the legacy system, the remaining 20% is the gnarliest, the cost-benefit on completing it is poor, and the project loses momentum. Six months later, the legacy system is still running, the new system is slightly bored, and the savings case has flatlined.
The right answer is to declare completion at 80% and decide what to do with the last 20% as a separate decision. Sometimes the last 20% is genuinely worth completing; sometimes it should be replaced with a different solution; sometimes it should be left alone because it’s a stable, paid-off system with no real reason to migrate.
A modernisation programme that doesn’t have a clear “done” criterion will run forever. The criterion has to be written down at the start, has to be reviewable, and has to be respected. We have retired everything that meets the cost-benefit threshold; the remainder is documented and scheduled separately. That sentence, on the page, is what closes the project.
A piece of work that closes is a piece of work that’s been finished. Most of the modernisation work we get called into hasn’t been finished by the people who started it. The pattern of the stalled 60% migration is the same pattern, in a different domain.
If a strangler-fig programme of yours has stalled, hello@altostack.io.