Migrations rarely fail because the technology did not work. They fail because something was not discovered beforehand, and the discovery happens at the worst possible moment — during cutover, with everyone watching.
These are the failures that actually cause downtime, roughly in order of how often they occur.
1. Undiscovered dependencies
The most common cause of a migration outage by a wide margin. An application is moved successfully and three other things stop working, because they depended on it in a way nobody had documented.
The usual culprits:
- Hard-coded server names or IP addresses inside applications and scripts
- Scheduled tasks running on a server nobody thought was important
- Reporting tools pointed at a database that moved
- Mail-enabled devices relaying through a local server — scanners, alarm systems, alerting tools
- Mapped drives referenced in documents, templates and shortcuts
- Integrations between line-of-business applications that nobody maintains
The prevention is a dependency map produced during discovery, and the honest way to build one is to ask what talks to each system rather than what each system talks to. The second question finds what the owner knows about. The first finds what they do not.
2. Bandwidth arithmetic nobody did
Two distinct failures here.
The first is the initial data transfer. Moving several terabytes over a business internet connection takes considerably longer than people estimate, and teams discover mid-migration that the copy will not finish inside the maintenance window. Do the arithmetic before committing to a date, and for large volumes consider a physical transfer appliance.
The second is ongoing. After migration, traffic that used to stay on the local network crosses the internet connection. A connection that was comfortable for email and web browsing can be inadequate for the same staff working entirely against cloud services, and the symptom is not an outage but everything being slow, which is harder to diagnose and easier to blame on the migration generally.
3. No rollback position
Teams commit to a cutover with no defined way back. When something goes wrong at 2am, the only options are to push forward through an unknown problem or to improvise a reversal that was never planned.
A rollback plan needs three things stated in advance: what triggers it, who decides, and how long it takes. Without the third, the decision gets made too late — you cannot choose to roll back at 5am if rolling back takes four hours and staff arrive at 8.
It also needs a verified backup taken immediately before cutover. Not a backup from the scheduled job the night before — a confirmed, tested backup at the point of no return.
4. DNS handled badly
DNS causes more visible migration problems than any other single technical component, and almost all of them are avoidable.
- TTL not lowered in advance, so a change that should propagate in minutes takes hours or a day
- Records missed — the ones nobody thinks about until something breaks, including autodiscover, SPF, DKIM and any record pointing at the old infrastructure
- Forgetting that mail will arrive at both old and new destinations during propagation, and not keeping the old system receiving
- Domain registrar credentials nobody has, discovered on cutover night
That last one is worth checking weeks in advance. Domain control is frequently held by whoever set it up years ago, sometimes a former employee or a previous provider.
5. Licensing discovered late
Software licensed for on-premise deployment may not permit cloud hosting, may require a different licence type, or may have no cloud equivalent at all. Vendors vary enormously in how they treat this, and some require a commercial conversation with a lead time.
Discovering this during migration planning is an inconvenience. Discovering it during cutover means either an unlicensed production system or a workload that has to move back.
6. Migrating the mess
Lifting an environment as-is carries every existing problem into the new one, where it is now costing you monthly rather than sitting on hardware you already own.
Specifically: stale directory accounts, over-provisioned servers sized for a peak that no longer exists, applications nobody uses, and permissions that accumulated over a decade. Discovery is the one moment when cleaning all of this up is easy, because you are looking at it anyway.
Over-provisioning deserves particular mention because it is the largest source of ongoing cloud waste. Servers sized like physical hardware — for peak load, permanently, with headroom for a five-year life — cost the same every month whether the capacity is used or not.
7. Cutting over everything at once
Big-bang migrations concentrate all the risk into a single window. If something fails, everything fails simultaneously and the team is diagnosing multiple unfamiliar problems under time pressure.
Waves are almost always better: one department, then a few more, then the rest. Each wave surfaces problems while the blast radius is small, and the team learns the process on people who can tolerate disruption before moving the ones who cannot.
Where a big-bang cutover is genuinely necessary — email, usually — pilot with a small group beforehand for long enough to find the problems that only appear in real use.
8. Declaring victory at cutover
The environment works, everyone moves on, and several things never happen:
- Backup is never configured in the new environment, because the old backup covered the old servers
- Security configuration is left at defaults
- Resources are never right-sized, so the monthly bill stays at the initial generous provisioning permanently
- The old environment is never decommissioned, so you pay for both indefinitely
- Documentation still describes the previous architecture
Schedule the post-migration review at the same time you schedule the cutover — 60 to 90 days out, with those five items on the agenda. It is the single cheapest way to avoid the most expensive long-term outcomes.