
Imagine the following scenario. Leadership at your company is pushing hard for data to be “AI-ready” and everyone is lining up to meet those goals. The objective is clear, but how do you get there in a way that works?
Although all AI projects are different, they also all run straight through the data. So then another question–is your data ready? Sadly, in most cases, your data isn’t AI-ready, and that simple problem is a big one. Because of it, only 7% of enterprises in one survey said their data is completely ready for AI.
Which raises the question: what does data need to be AI-ready?
One basic thing is data access. Without access, you can’t even begin to use your data, so just as all roads to AI lead through data, all roads to data lead through access. But how do you get that access, given that data is spread across multiple data sources? The solution to this calls for universal data access, thereby giving your AI agents the ability to find and work with data across your organization. In this sense, AI has made universal access a necessity.
Once upon a time, the only technically feasible approach to universal data access was data centralization, involving migrating all of your data into one data source. But this doesn’t scale in the AI era.
Fortunately, there’s a better way to approach data access, and that involves data federation. In this article, we’ll look at how centralization came to be the perceived wisdom in the data industry, how data federation provides a better path forward for AI agents, and how (and when) centralization still fits into the picture.
What it means for data to be AI-ready
Universal data access is indispensable for AI-readiness. But there’s more to it than that.
AI agents need large amounts of relevant business context in order to produce accurate, relevant answers that avoid hallucination. The large language models (LLMs) that power AI can process human language and general facts. They know nothing, by default, about how your business works or what your customers expect from you.
What’s more, much of this data won’t be AI-ready by default. An AI agent may have no idea what customer_value_003 means. Two data warehouses may define quarterly sales differently, leaving an agent unsure which one to use and when.
Your data also likely contains large gaps in its context. Humans do a lot of work using intuition and institutional knowledge. For example, your AI agent doesn’t know what customer_value_003 means, but one particular human in Accounts Payable certainly does.
Accessing this context, associating it with existing data, and making this context available to AI agents across the enterprise is what’s called context engineering. This entails:
- Getting definitions of core metrics out of the realm of institution and into a concrete semantic layer, so “quarterly sales” has one distinct, agreed definition.
- Using data products to codify this logic and connect data sources to AI using trustworthy packages of data that contain the metadata and meaning AI agents need to engage in proper decision-making.
- Recording data lineage so an AI agent knows where data comes from and whether it can be trusted.
- Centralizing data governance so that an AI agent’s access to data can be regulated and audited in one place, not across dozens of disparate data systems.
The end result of this work is a context layer. The context layer sits above your storage and compute, providing the meaning and governance your AI agents would otherwise lack.
Why centralization by default isn’t the right choice for AI
More companies are realizing the importance of context engineering. The instinct is to implement it with centralization. That’s how most of us, after all, attempted to solve similar challenges with data analytics and AI.
At first glance, centralizing everything seemed like a reasonable solution. Putting all your data in one place solved your data access problem by definition. It also simplified discovery, collaboration, and governance. What better way to access than to put it all in one place?
In practice, centralization by default always ran headlong into several brick walls, even for data analytics. The approach proved:
- Time-consuming. Moving, rationalizing, and transforming terabytes of data into a single location could take weeks, often months. During that time, all data-dependent projects become frozen until the massive centralization project is finished.
- Costly. The compute, storage, and person-hours required to move everything into one place made most centralization projects prohibitively expensive. Much of this data ended up going unused, meaning that the effort put into centralizing it was wasted.
- Unending. The growth rate of data isn’t slowing down, but speeding up. Previous figures pegged the annual global data growth rate at around 23%; IDC predicts it’s growing even faster now, thanks to AI. A centralization project, therefore, can never be finished. You’ll constantly be moving more data into an increasingly unwieldy repository.
- Rigid. Rather than keeping data in heterogeneous systems and using the best tool for the job, you moved it into one vendor’s data warehouse or data lake solution.
AI compounds each of these problems. A data analytics workload required merging data from a half-dozen or so systems. Agentic solutions require high-quality, contextualized data from every corner of your company. The scale of the effort makes centralization by default a non-starter.
Centralization 2.0: Federate, then centralize selectively
Does centralization still have a part? Absolutely. Centralization can be the right call, as long as it’s done selectively. The solution is to flip the problem on its head. Instead of centralizing everything, use data federation coupled with a high-performance data lakehouse for selective centralization.
Data federation connects your AI to your data immediately, not weeks or months from now. With data federation, you can embrace a hybrid architecture that enables both decentralized and centralized data scenarios, with centralization happening on your terms, in line with your business needs.
Today, we have the tools for any company to implement a distributed, performance-first architecture. Distributed query engines like Trino, around which Starburst is built, combined with modern open table formats like Apache Iceberg, let us work with most data where it lives while moving critical datasets into the lakehouse, where they benefit from Iceberg features such as reduced metastore reliance, optimistic concurrency, and managed table maintenance with LakeOps, among many others.
Centralization was once a necessity. Now, it’s an option. Data federation with a data lakehouse enables an iterative, DevOps-style approach to centralization that fits within your overall AI context engineering efforts:
- Start small. Use data federation to integrate your existing data where it lives today into your context layer, where you can add metadata and create consistent metrics.
- Identify your highest-value use cases. Monitor data usage and performance to find the raw data sources that could benefit from centralization.
- Centralize selectively. Move high-value data closer to your agents by migrating it into your lakehouse.
Creating a distributed context layer for AI
Data centralization didn’t come out of nowhere. At the time, it was the best solution we had to the hard problem of accessing trustworthy, consistent data.
Since then, the demand for data has grown astronomically, driven most recently by AI. Fortunately, the tools we have for managing data have grown with it.
Building that context layer requires the right platform. Starburst is built from the ground up to support data federation and universal data access. It supports context engineering through the creation of curated data products that are consistent, discoverable, and versioned. Governance controls, such as role-based access control, travel with your data products, so your BI tools, APIs, and AI agents can all access the same definitions safely.
Starburst streamlines data access for agents with our hosted MCP server, so that agents can navigate data without hard-coded APIs. Analysts and business stakeholders can access the same data products directly using AIDA, our conversational agent that turns natural-language questions into answers.
If you’re interested in making the move from data centralization to an AI-ready distributed context layer, talk to us – we’re here to help.



