An AI data strategy is a plan for using AI to improve how you find, classify, govern, and prioritize your data, not a wish list of AI projects waiting for budget. This guide walks through the eight steps we use to get there, in the order that works. It is written for data leaders in European enterprises: CDOs, CIOs, heads of data, and the architects who live with the decisions. Eurostat reported that 19.95% of EU enterprises used AI technologies in 2025, up from 13.5% in 2024, and adoption moving at that pace tends to outrun the governance work underneath it.
Four decisions make up an AI data strategy: which data the business needs, who owns it, how it is governed, and which parts of that work AI can accelerate. It sits above tooling and below business strategy, and none of it requires a new platform purchase.
Two things trip teams up before step one. Access comes first. AI-assisted discovery needs read rights on metadata, a narrower permission than production data access and usually faster to approve. Authority comes second. Someone has to be able to reassign ownership across business domains, and that person is rarely in the data team.
| What you need | What that means in practice |
|---|---|
| Executive sponsor | A named CDO, CIO, or COO who can approve ownership changes across at least two business domains |
| Source system list | Every database behind every in-scope application, with engine type, version, and hosting model |
| Metadata read access | Catalog or database rights to schemas, table names, and column comments, without production data extraction |
| Data protection contact | A named DPO or counsel who can rule on lawful basis for reusing personal data |
| Domain owners | One business owner for each of two or three domains, for example billing, customer, or network |
| Written outcomes | Three to five outcomes stated as targets, for example cutting monthly close from 12 days to 5 |
| A metadata tool with an owner | Unity Catalog, Microsoft Purview, Snowflake Horizon, or equivalent, already administered by someone |
Work through these in order. Steps 2 through 4 are where AI does the heavy lifting, and steps 5 through 8 are where humans commit to decisions.
Write three to five outcomes in business language, each with a metric and a named owner. Cost reduction, faster regulatory reporting, fraud detection, personalization, and data monetization are all valid starting points. If an outcome cannot be measured, it is a theme. You are done when every line on that page has a number attached and a person accountable for it.
Point AI at metadata rather than at data. Follow Microsoft’s workload assessment guidance, which says teams should inventory every database an application uses. That means engine type, version, and hosting model, plus whether each one is self-hosted, virtual-machine hosted, or a managed service. Done properly, this leaves one catalog entry per source system, with an owner field that may still be blank.
Run two classifications at once, because a table can be highly sensitive and commercially worthless. AI proposes labels from column names, sampled patterns, and lineage, and a human confirms them. In Europe the sensitivity axis has to include GDPR personal data, regulated-sector data, data used to train models, and data generated by connected products under the European Commission’s Data Act. The output to aim for is every in-scope table carrying one sensitivity label and one business-value tier.
Ask the model to draft quality rules per domain, flag anomalies, group schema drift by source system, and summarize where lineage breaks. This is drafting work, not deciding work. Every rule needs a human to set the threshold that counts as fit for purpose, because acceptable completeness in a marketing table is unacceptable in a regulatory report. Finished, this produces a ranked defect list per domain, each item with an owner and a threshold.
Move from “all data matters” to a small number of domain data products. Each needs an owner, known consumers, a service level, stated quality expectations, and a business use you can point at. Two or three is the right number to start. The test is whether a consuming team could build against the spec without asking follow-up questions.
Do not stand up a second committee. Data governance and AI governance share the same evidence base, which is lineage, classification, access, and documentation. The EU AI Act entered into force on 1 August 2024, and its transparency obligations and AI literacy duty apply from 2 August 2026. The Digital Omnibus on AI then entered into force on 27 July 2026, deferring the stand-alone high-risk obligations to 2 December 2027 and those for AI embedded in regulated products to 2 August 2028, according to the European Commission. Treat that as preparation time rather than a reprieve, because the underlying requirements did not change. ISO/IEC 42001, which ISO describes as the world’s first AI management system standard, gives the management-system frame for the same work. What you want at the end is one register listing data domains and AI systems together.
Now compare warehouse, lakehouse, data fabric, data mesh, and hybrid patterns. The choice follows from what steps 5 and 6 produced: how many domains own their data, how strict residency requirements are, and how mature the platform team is. Eurostat reported that 52.74% of EU enterprises used paid cloud computing services in 2025, rising to 84.67% among large enterprises. Write the outcome up as a short architecture decision record naming the pattern, the reason, and what was rejected.
A strategy nobody can start is a document. Sequence the first 90 days across quick wins, governance foundations, the data products, the platform decision, and the operating-model changes that make ownership stick. Put dates and names on everything. Success looks like a roadmap whose first item can begin next Monday without another approval cycle.
In Europe the governance work is part of the design, not a compliance appendix. A US-headquartered playbook can treat classification as a security exercise, while a European estate has four overlapping regimes shaping the same decision.
Start with the AI Act. Its risk-classification logic means the answer to “can we build this use case” depends on what the system does and who it affects, which has to be decided before engineering starts rather than after a pilot succeeds. The high-risk deferral to December 2027 moved the deadline, not the work. Then add the Data Governance Act, which the European Commission describes as a pillar of the European data strategy aimed at increasing trust in data sharing and removing technical obstacles to reuse. That reframes internal data sharing as something to enable deliberately, not restrict by default.
The practical consequence is sequencing. Classification and ownership move earlier in a European strategy than they do elsewhere, because both the lawful basis for reuse and the AI risk classification depend on them. Teams that leave governance until after the platform build end up re-cataloging the same estate twice.
The useful split is between drafting and deciding. AI produces first drafts at a speed no team can match, and it has no standing to accept the risk attached to any of them.
| Strategy task | What AI does well | What a person has to decide |
|---|---|---|
| Data estate inventory | Extracts metadata, clusters tables into candidate domains, flags duplicate columns across systems | Which domains are in scope and who owns each one |
| Sensitivity classification | Proposes labels from column names, sampled patterns, and lineage paths | Whether the label is correct under GDPR and what the lawful basis for reuse is |
| Quality assessment | Drafts rules per domain, detects anomalies, groups schema drift by source | The threshold that counts as fit for purpose in that domain |
| Lineage documentation | Summarizes where lineage breaks and writes first-draft field descriptions | Whether an undocumented pipeline gets retired, rebuilt, or accepted as is |
| Use-case prioritization | Ranks candidates against data availability and build effort | The AI Act risk classification and the go or no-go |
| Roadmap drafting | Produces a first-pass sequence from dependency data | Budget, timing against business cycles, and who is accountable |
Read the right-hand column as a workload, because it is one. Every row AI accelerates creates a decision queue, and strategies stall when that queue has no owner, not when the tooling underperforms.
Take a mid-sized European operator with billing, CRM, network, and digital event data in four separate systems and a mandate to launch AI-assisted customer service. In week one the team writes four outcomes. In weeks two and three, AI-assisted discovery produces a catalog of source systems and a proposed domain map, which the team corrects in a half-day workshop. Classification runs next, and the DPO rules on two contested tables. By week six there are two data products with owners, and only then does the architecture conversation start.
Here is the part that surprises people. In the data platform work we have run at Exacaster since 2011, the step that fails most often is not the platform choice. It is that the domain has no real owner once the pipelines go live, so quality rules quietly stop being enforced and the catalog drifts out of date within two quarters. The fix is unglamorous. Put the owner’s name in the roadmap before you put the platform’s name in it.
That ordering protects the AI work downstream. An assistant built on a domain with a named owner and enforced quality rules can be audited when someone asks why it answered as it did. One built on an unowned dataset cannot.
Five mistakes recur often enough to be worth naming.
Treating the catalog as the strategy. An AI-generated inventory is an input. Until it has owners, thresholds, and dates attached, no decision has been made.
Classifying only personal data. GDPR is the familiar axis, so teams stop there and miss connected-product data under the Data Act and data used to train models. Both carry obligations that surface later, usually at the worst moment.
Letting AI set quality thresholds. Models draft plausible rules. A completeness target that suits a campaign table will fail an audit on a regulatory report, and only someone who owns the domain can tell the difference.
Choosing architecture first. Picking a lakehouse before you know how many domains will own their own data means the architecture decides the operating model by accident.
Running two governance tracks. Separate data and AI committees produce two registers, two sets of documentation, and contradictory answers about the same dataset.
Do the first two steps internally, always. Nobody outside the business can write its outcomes. External help earns its cost around step 4 or step 6, when quality and lineage work stops being a project and becomes an operating burden, or when AI governance has to satisfy an auditor rather than a steering committee. If that is where you are, we at Exacaster help with data strategy, platform builds, and managed data operations. We also work on moving AI use cases from pilot into production through our generative AI accelerator.
What is the difference between a data strategy and an AI strategy?
A data strategy decides which data the business needs, who owns it, and how it is governed. An AI strategy decides which use cases to build and how to control their risk. The AI strategy depends on the data strategy, and it fails without it.
Can AI automate data strategy from end to end?
No. AI can automate discovery, documentation, and drafting, roughly the mechanical half of the work. It cannot accept regulatory risk, assign ownership across business units, or commit budget, and each of those is a strategy decision.
How long should the first version take?
Six to ten weeks for a scoped estate, ending in a 90-day roadmap. Teams aiming for full coverage before deciding anything typically spend two quarters producing a catalog nobody acts on.
Does the EU AI Act apply to internal-only tools?
It can. The obligations follow the risk the system creates and who it affects, not whether it is customer-facing. An internal tool that influences decisions about people is worth classifying before it is built.
Where should you start if the estate is very large?
Pick two or three revenue-critical domains and finish them completely, including owners, quality thresholds, and one data product each. A finished narrow slice teaches the organization more than a wide inventory with nothing decided.
That gives you an order of operations: outcomes, inventory, classification, quality, data products, governance, architecture, and roadmap. The sequence matters more than any single step, because most expensive rework in European data programs comes from deciding architecture before ownership. Start with the one-page outcomes list this week, and talk to our data team if you want a second opinion once the roadmap is drafted.