Gradient blur
Je AI-strategie staat of valt met je datafundament
Blog
Artificial Intelligence
Data & Analytics
Managed Services
Back to overview

Your AI strategy stands or falls with your data foundation

Last year I got a phone call that I've since heard come back in all kinds of variations. "Peter, we've done everything that seemed logical.
21 - 04 - 2026

Azure OpenAI, Copilot Studio, the right guidelines followed. And yet the output stays subpar. Where does it go wrong?" After half an hour of digging, it was clear: the problem wasn't in the model, but in the data. Outdated SharePoint documents, ERP exports in Excel, and a knowledge base with several definitions sitting side by side. Everything was there. It just didn't really work together. The model did what it was supposed to do. The data didn't. As Data & Analytics Business Lead at Xylos, I see this pattern in almost every organization that wants to get started with AI. That's why this article is about the data layer: the foundation that gets far less attention than the model, but determines far more.

This is the third article in our blog series about the building blocks of a scalable AI approach in organizations. First we looked at AI sprawl in organizations. Then at the role of the Power Platform as a bridge to controlled AI integration. This time it's about the foundation under that approach: the data layer.

The model is rarely the bottleneck

An intelligent system is only as smart as the data it runs on. Garbage in, garbage out, only now at scale and with a hefty price tag.

GPT-4, Llama, Mistral, Phi-3: for most enterprise use cases, today's models are more than strong enough. The difference between an AI solution that builds trust and one that causes frustration usually lies elsewhere.

The real bottleneck sits in what happens before the model. Is the data available? Is it up to date? Does everyone understand the same definitions? Does anyone know who owns the source the output relies on?

Take a retail organization that wants to build an AI assistant for demand forecasting. The model is ready within a few weeks. The integration between ERP, WMS and external market data drags on for seven months, because no one ever decided how those flows come together, who manages them, and which version is reliable.

That's not an exception. That's the pattern.

An intelligent system draws its value from the quality of the context it operates on. As soon as that context becomes fragmented, outdated or unclear, the error scales right along with it. Only now faster, more convincing, and more expensive.

What a data layer for AI really needs to do

A data layer isn't a storage space where you gather everything and then hope AI does something meaningful with it. It has to deliver on four things at once.

1. Availability. Relevant data must be accessible through a logical, coherent layer, regardless of where it historically originated.

2. Quality. Data must be validated, documented and managed. Without ownership, quality becomes a matter of chance.

3. Freshness. The data has to match the speed your use case requires. For some scenarios, a daily update is enough. For others, near real-time is needed.

4. Governance. You need to know who has access to which data, why, and whether that aligns with your policy and compliance obligations.

On paper, that sounds obvious. In practice, we see that in most organizations at least one of those four is under pressure. Often more than one at a time.

Why Fabric shifts the conversation

Microsoft Fabric changes that conversation because it starts from integration rather than fragmentation. Where data used to be spread across separate tools for integration, engineering, warehousing, BI and governance, Fabric brings those layers together within one platform, with OneLake as the shared foundation.

That has a direct impact on AI. Once your data no longer has to pass through separate exports, duplicate storage and improvised connections, the reliability of what an AI system retrieves and generates goes up. You reduce latency, limit data pollution, and make governance structural instead of bolted on afterwards.

For use cases that need to respond faster, Fabric also adds native capabilities around real-time intelligence. And with Purview as a built-in layer for policies, compliance and auditability, visibility finally becomes part of the foundation.

That nuance still matters: Fabric is an enabler, not a miracle cure. Organizations only really benefit from it once they've first made clear what their data model looks like, which domains exist, and who's responsible for what. Fabric makes a good strategy executable. It doesn't replace one.

Why the lakehouse is becoming the logical AI architecture

The shift toward lakehouse architecture isn't hype. It follows directly from what AI demands of data.

A classic data warehouse excels at structured, historical analysis. That remains valuable for reporting and management information. But it hits its limits faster once AI also needs to take in text, documents, images or event data. A data lake handles that volume story better, but without enough structure it risks degenerating into an environment where everything looks available and no one knows anymore what's reliable.

The lakehouse brings those two worlds together. The flexibility of a lake. The reliability and manageability of a warehouse. That exact combination is what makes it suitable as the foundation for modern AI workloads, where structured and unstructured data together form the context for agents, copilots and analytical applications.

Three concrete recommendations for your organization

1. Start with a data audit
Not to file an inventory away in a drawer, but to get a sharp view of what data you actually have, where it lives, how current it is, and who manages it. That almost always uncovers pain points that were previously invisible.

2. Define your data strategy first
Choosing technology without clarity on domains, definitions and ownership rarely leads to acceleration. Usually, you end up building faster on a foundation that isn't stable enough yet.

3. Treat governance as a growth accelerator
Governance is still too often seen as a brake or an obligation. In reality, it lets AI use cases go live faster, precisely because trust, access and control are arranged in advance.

The organizations that will make the difference with AI in the coming years aren't necessarily the ones that are first to experiment with a new model. They're the organizations that get their data foundation in order today.

The real question for the next phase

That's also the core of this story: an agent with access to the wrong context stays just as unreliable, no matter how smart the model seems.

At Xylos, we help organizations get exactly that right: from data assessment to lakehouse architecture and Fabric implementation, always starting from the business value a solution needs to deliver. Get in touch for a no-obligation data assessment. We'll map out where you stand today, where the biggest gaps are, and which steps deliver the most return.

In the next article, we go one step further. We'll look at how you connect this data foundation to AI output you can actually trust and verify: grounded AI, and the architecture needed to make that happen.

About the author

Peter Verrykt is Data & Analytics Business Lead at Xylos and guides organizations in turning data into concrete business value. He helps companies look beyond technical implementations and use data as a foundation for better decisions, greater agility and sustainable growth.