There is a recurring shape to the work we get called into. A migration started two or three years ago. Forty to sixty percent of workloads have moved. The remaining workloads — the ones the original programme didn’t reach — are now the most consequential ones in the estate. The original team has either disbanded, moved on, or been absorbed into BAU. The cloud bill is up. The on-premises footprint is also up, somehow. Someone is asking, sensibly, whether the programme should be declared finished and the rest left where it is.
It usually shouldn’t. But the reasons it shouldn’t are not the reasons that get the programme restarted. This is a post about the shape of stalled migrations, why finishing them is more expensive than the original budget said, and what the second attempt has to do differently to land.
Why migrations stop at 60%
The first 60% is the workloads that fit the playbook. Web tiers behind load balancers. Stateless services that scale horizontally. Batch jobs that read from one source and write to another. Anything that survived a containerisation effort or was already on virtual machines. The playbook works because the playbook was written for these workloads.
The remaining 40% is the workloads that don’t fit the playbook. Not because they’re unusually complex — though some are — but because each one has at least one property that the playbook didn’t anticipate. A licence the vendor refuses to relicense in cloud. A storage performance characteristic that’s plausible on-premises and ruinous in cloud. A network adjacency to a system that hasn’t moved and isn’t going to. A regulator who hasn’t yet been consulted on whether moving this workload is allowed.
The first 60% is the work the playbook was written for. The last 40% is the work the playbook was written to avoid.
The original programme made the rational call to ship the easy 60% first. Then, having shipped it, the programme ran out of momentum at exactly the point where the work stopped being a playbook and started being individual judgement calls. The team that excelled at industrialising the easy work was, by composition and by tooling, the wrong team to do the hard work. The hard work was deferred. Two years later, deferred is structural.
Shape one: the workload nobody owns
The most common stall. A workload sits in a corner of the estate, runs reliably, generates revenue, and has no engineer whose job description includes it. The original owner left. The team rotated. The system is now “owned” by a wiki page with three names on it, two of whom no longer work at the company.
The migration stalled here because nobody could answer the basic questions. What does this depend on? What depends on it? What’s its tolerance for downtime? What’s its data residency requirement? The original migration team escalated, got vague answers, decided this was not the workload to die on a hill for, and skipped it.
Finishing this one costs more than the original budget assumed because the first month is forensic. Reading the system is an engineering project in its own right — sometimes ten engineering days, sometimes thirty, before any migration work has happened. The price tag is real and unavoidable. The compensating saving is that workloads in this state are usually smaller than they look in the inventory, because half of what they appear to depend on turns out to be unused.
The mistake to avoid: don’t try to migrate this workload into a clean cloud-native architecture on the first pass. Re-host first, in a configuration as close to the existing one as the cloud will allow. Modernise after, when the team has read the system properly. Trying to do both at once is how this workload stalled the first time.
Shape two: the network adjacency
A workload that talks to twelve other systems, eleven of which are on-premises, is not a candidate for migration on its own. It’s a candidate for moving as part of a wave that includes the systems it talks to most, with the rest reachable over a controlled cross-environment path.
When the original programme stalled, it was usually because the wave the team designed turned out to depend on workloads that the network team, the security team, or another business unit owned and had not agreed to move. The dependencies surfaced in week six of an eight-week wave plan. The wave was rescoped. The workload didn’t move.
Finishing this one costs more because, the second time around, the wave shape has to be redesigned around what’s actually still available to move. Often that means the workload moves with most of its dependencies still on-prem, reached through a more elaborate networking layer than the original architecture wanted. AWS Direct Connect or Azure ExpressRoute for the steady-state traffic; Transit Gateway or Virtual WAN for the routing; private DNS arrangements that have to be agreed with three teams.
That’s a real engineering project. It’s also the only honest path. Pretending the dependencies are going to move on the second attempt — when they didn’t on the first attempt and nothing has changed — repeats the stall.
Shape three: the data tier the source-system owners won’t sign off on
A migration that depends on moving a database the business considers regulated, sensitive, or revenue-critical will not move until the people responsible for that database have signed off on the destination architecture. Sometimes those people are a regulatory function. Sometimes they are a senior engineer who, fairly, doesn’t trust the cloud target you’ve designed.
The original programme stalled here because the migration team approached the conversation as a technical one. They presented the architecture. They answered the technical questions. They didn’t get the sign-off, because the question wasn’t really technical. The question was “do I trust the team that’s going to operate this in cloud, given that I currently operate it on-premises and know exactly what happens at 03:00”.
Finishing this requires a different conversation. Not “look at our architecture”, but “look at our operations”. Show the runbook. Show the alert routing. Show the rota. Show the most recent incident review. Demonstrate, with operational artefacts, that the system in cloud will be operated to the same standard or better than it currently is. The signoff usually follows once the operational evidence does.
The price of this work is mostly time and visibility — not engineering effort in the migration itself. A migration team that hasn’t budgeted for it will run out of runway in the conversation, not in the build.
Shape four: the licence that costs more in cloud than on-premises
A subset of legacy workloads run on software whose licence terms make the cloud version uneconomic. The original programme either didn’t notice or noticed and didn’t have authority to renegotiate. The workload stayed where it was.
This one resolves three ways. First, replace the licensed component with an open-source or cloud-native alternative — a real engineering project, with the risk of behavioural drift, but often a positive ROI within twelve months and the path that resolves it permanently. Second, run the licensed component in a hybrid configuration: the bulk of the workload moves; the licensed piece runs on dedicated infrastructure, sometimes literally on the on-premises host with a Direct Connect into the cloud workload. Third, renegotiate the licence — usually the cheapest path if the vendor is willing, often the most political.
The mistake the second attempt has to avoid is treating this as a procurement problem and waiting for procurement to resolve it. Procurement timelines for licence renegotiations are unpredictable and often slow. The technical option — replacing the component — has to be running in parallel as the contingency. Otherwise the migration is hostage to a contract negotiation it has no leverage in.
What finishing actually costs
The honest number for completing a stalled migration is roughly 1.5x to 2x what the original budget said the entire migration would cost. Not 1.5x of the remaining 40% — 1.5x of the original whole.
The components of that number are predictable.
Forensic engineering on shape-one workloads, before any migration work, runs 10-30 engineering days per workload. With a long tail of unowned workloads, this compounds quickly.
Network redesign for shape-two workloads is usually a calendar-quarter project of its own, with three-team sign-offs, before the migration build can resume.
Operational evidence-gathering for shape-three workloads is often a full quarter of running the existing system to the new operational standard — building the runbooks, the SLOs, the incident review cadence — so that the cloud target is a like-for-like, not an upgrade-and-pray.
Licence work for shape-four workloads is contingent: cheap if the vendor cooperates, expensive if the replacement engineering has to land first.
On top of that, every stalled migration we’ve recovered has had a hidden cost the original budget didn’t account for: the cost of running both estates at once. The on-premises footprint hasn’t shrunk on the schedule the cloud bill grew on. For two years, the company has been paying for the floor space, the power, the maintenance contracts, and the cloud spend. That bill is what makes the second attempt economically obvious, even at 1.5x. The third year of dual-running is more expensive than finishing.
What the second attempt does differently
A second attempt that lands has four properties the first one didn’t.
It has a written done-criterion. We have moved everything that meets the cost-benefit threshold; the remainder is documented and scheduled separately. That sentence, agreed at the start, is what closes the project. Without it, the second attempt becomes the third attempt by month nine.
It has a different team shape. The original migration team excelled at the first 60%. The second attempt needs people who excel at the last 40% — engineers who’ve read estates as old as this one, who treat unowned workloads as a forensic challenge rather than a process violation, who know the network and licence patterns by name. That team is smaller, more senior, and more expensive per head, and it ships less code per week. It also doesn’t stall.
It runs as a sequence of small waves with rollback at every step, not as a multi-quarter programme. Two-week iterations, each with a single workload or coherent group, demoed live, with a defined rollback path. Production from day one — no staging-only theatre. The same finishing standard the original programme should have been held to held to now.
And it has the operational hand-off built in from the start. Day-2 runs from the migration team until the receiving team is ready. SLOs, alert routes, incident review, postmortems shared in the receiving team’s tooling. The migration is finished when the rota tells you it is.
The trap of “good enough”
Every stalled migration has a tempting alternative: declare it finished, leave the remaining 40% on-premises permanently, accept the dual-running bill as a steady-state cost. Sometimes that’s the right answer — for a small, declining workload with a five-year retirement horizon, it can absolutely be the right answer.
For most stalled migrations we read, it isn’t. The “good enough” calculation usually ignores the compounding cost: the on-premises tail isn’t only paying its own bills, it’s anchoring the company to a regulatory regime, an operating model, and a hiring market that gets harder every year. The cost of leaving the tail isn’t this year’s bill. It’s the inability to shut down the data centre, the inability to consolidate operations, and the inability to put the migration on the “done” list of the organisation.
Finishing isn’t a virtue in itself. Finishing is the thing that makes everything next cheaper. The migration that’s 60% done is more expensive than the migration that’s 100% done, every year, until you finish it.
Stalled migrations are most of what we recover.