Cloud data migration is the process of moving enterprise databases, pipelines, and analytics workloads from legacy or on-prem systems into cloud platforms, done in a way that keeps reporting running and stays inside European data rules. This guide is written for data leaders, architects, and transformation teams who own the migration and have to answer for both uptime and compliance. We will cover what to confirm before you start, the ten steps in order, how to choose a migration pattern, how residency shapes the plan, and the mistakes that cause reruns. The aim is a migration you can sequence, test, and reverse, not a single risky cutover.
A cloud data migration moves data and its pipelines into cloud infrastructure while preserving accuracy, access, and compliance. Before step one, a few things must be in place, or the project stalls at the first residency or reconciliation question.
| Prerequisite | Why it matters | Ready when |
|---|---|---|
| A named business driver | Cost, AI readiness, and end-of-life each imply different scope | You can name the reason and what success looks like |
| Workload and data inventory | You cannot migrate or test what you have not mapped | Every source, pipeline, and consumer is listed |
| Classification and residency map | EU rules depend on data type and location | Each dataset is tagged by sensitivity and allowed region |
| Target platform shortlist | Architecture and sequencing depend on the destination | A warehouse, lakehouse, or hybrid direction is chosen |
| Landing zone owner | Security and networking need one accountable owner | Identity, network, and access design has an owner |
| Rollback and reconciliation plan | Live reporting must survive each wave | A reversible plan and match criteria exist per wave |
The ten steps below move from business intent to a running, monitored platform. Each step names what a correct result looks like, so you know when to move to the next.
State the driver: AI readiness, analytics performance, cost control, regulatory need, data residency, end-of-life systems, or business agility. A correct result is a one-line reason that the scope and pattern choices can be checked against later.
List every database, pipeline, and downstream report. Microsoft’s workload assessment guidance recommends recording engine type, version, hosting model, and whether each database is self-hosted, VM-hosted, or managed. A correct result is an inventory where no critical dashboard traces back to an unknown source.
Tag each dataset by sensitivity and the rules that apply, including GDPR, DORA for financial entities, NIS2 for critical sectors, and the Data Act. A correct result is a map that tells you, per dataset, which regions and controls are allowed before any architecture is drawn.
Decide per workload whether to rehost, replatform, refactor, rebuild, replace, retire, or retain. For data specifically, also decide between batch migration, streaming replication, and phased domain migration. A correct result is a pattern assigned to each workload with a reason, covered in the decision section below.
Choose between a cloud data warehouse, a lakehouse, a data lake, a data mesh, or a hybrid design, and match the platform to the workload. A correct result is an architecture where the platform choice, whether a warehouse or lakehouse type, follows the workload’s query and volume needs rather than fashion.
Set up identity, networking, encryption, access policies, cataloging, lineage, monitoring, logging, and cost controls before bulk data moves. A correct result is a landing zone where access and lineage exist on day one, so the platform is governed rather than retrofitted later.
Define transfer paths, test criteria, and a rollback path for each wave. The Microsoft Cloud Adoption Framework migration planning guidance calls for workload sequencing, data transfer paths, rollback strategies, success criteria, and review checkpoints. A correct result is a wave plan where any wave can be reversed without data loss.
Pick a real but contained business domain, never a toy dataset, and validate performance, reconciliation, data quality, security, lineage, and user acceptance. A correct result is a pilot where the numbers match the source and the users who rely on the data accept it.
Group workloads by business domain, dependency, criticality, risk, and value, then migrate in sequence. A correct result is each wave signed off on reconciliation before the next begins, so problems stay contained.
Add observability, cost management, DataOps, incident response, and data-quality checks, then hand over to an operating model. A correct result is a platform with an owner, monitoring, and a cost view, not a migration that ends at go-live.
The pattern decision affects cost, timeline, and risk more than the platform brand does, and it is rarely the same answer twice across an estate. A stable finance database and a tangled reporting warehouse should not get the same treatment. That is why the choice belongs per workload, not per company.
| Pattern | Use when | Watch out for |
|---|---|---|
| Rehost (lift and shift) | The workload is stable and speed matters | Old quality and structure problems move with it |
| Replatform | You want cloud benefits without a redesign | Partial changes can leave awkward seams |
| Refactor or rebuild | The data model blocks analytics or AI | Highest effort and longest timeline |
| Retire or replace | The system is redundant or better bought | Hidden dependencies surface late |
| Retain (hybrid) | Residency or latency keeps data in place | Two environments to govern and secure at once |
Alongside the pattern, choose how data actually moves. Batch migration suits stable historical data, streaming replication suits systems that cannot pause, and phased domain migration suits large estates where you move one business area at a time. The practical rule is to rehost what is stable, rebuild what is blocking you, and use streaming replication wherever a business cannot tolerate a freeze.
In Europe, residency and compliance are part of the migration design, not a final review. GDPR governs personal data, DORA adds operational resilience duties for financial entities, and NIS2 raises security obligations for critical sectors. This is why classification sits at step three, before architecture, so a dataset is never designed into a region it is not allowed to occupy.
The AI angle sharpens the point. The EU AI Act entered into force on 1 August 2024 and becomes generally applicable from 2 August 2026, according to the European Commission. If your migration driver is AI readiness, the lineage, access, and documentation you build now become the evidence base your future AI governance depends on.
Most migration reruns trace back to a short list of avoidable errors. Naming them early is cheaper than fixing them mid-wave.
First, skipping classification. Teams that pick a region before mapping data sensitivity often find a dataset cannot legally sit where they put it, then redesign late.
Second, one pattern for everything. A blanket lift and shift carries old problems forward, while a blanket rebuild wastes effort on workloads that were fine.
Third, piloting on a toy dataset. A clean sample hides the reconciliation, volume, and permission problems that only real data reveals. Pilot on a real domain.
Fourth, treating “transfer complete” as success. A finished copy is not a verified one. Sign off on row counts, totals, lineage, and user acceptance, not a status bar.
Fifth, leaving governance for later. A landing zone without catalog, lineage, and access rules decays into the same mess you left behind. Build governance in from day one.
Running a migration in-house works when the estate is contained, the team has cloud and governance experience, and you can tolerate some downtime. It stops working when several pressures stack up at once. A large legacy estate, strict European residency rules, reporting that cannot break, and a fixed deadline together are a strong signal to get help. If your project is at that point, we at Exacaster help enterprise teams plan and run migrations across cloud and on-prem environments, and we tend to add the most value at the classification, pilot, and reconciliation stages. If broader modernization rather than a straight migration is your real goal, a companion guide on data modernization in Europe is a useful next read.
How do you reduce risk in a cloud data migration?
Migrate in reversible waves, pilot on one real business domain, and make reconciliation the sign-off gate. Risk drops sharply when each wave can be rolled back and when success is measured by matched data, not a completed transfer.
How do you test that a migration succeeded?
Compare row counts, key totals, and aggregates between source and target, confirm lineage is intact, and get sign-off from the people who use the reports. A migration is verified when the numbers reconcile and users accept the output.
What is the difference between data migration and data modernization?
Migration moves data and workloads to a new location, often as they are. Modernization also improves quality, governance, and structure. You can complete a migration and still need modernization if the underlying data model stays messy.
Do we have to move all data at once?
No, and you usually should not. Wave-based migration grouped by business domain contains risk and keeps reporting stable. Streaming replication lets systems that cannot pause stay live while data moves in the background.
Where do migration budgets most often overrun?
On underestimated source-system sprawl and data quality cleanup, not on the target platform license. Undocumented pipelines and duplicate data appear late, which is why the inventory and classification steps repay the time spent on them.
You now have the order that keeps an enterprise migration steady: define the driver, inventory and classify, pick a pattern per workload, design the target, build a governed landing zone, then pilot, migrate in reversible waves, and operate. The most useful next step is to draft your classification and residency map and pick one real domain for a pilot, because those two decisions shape the whole sequence. If you would like a second view on that plan before committing to a wave schedule, we are happy to talk it through.