
Something haunts AI success, and it’s hiding in plain sight. While most people are focused on the most overt aspects of the AI data stack–the models, the chipsets–trouble is brewing further down, in the subterranean foundation underneath the stack itself.
I mean, of course, the data.
The AI story is, at its core, one about data. It’s all about the data. Follow the data, and you follow AI back to its roots and the foundations that sustain it and drive it forward. AI feeds on data. It requires it. But getting at that data isn’t easy, and more than that, the way most organizations are approaching it is incorrect.
The reasons for this are both important to this revolutionary moment right now–the AI era–a modern phenomenon sweeping us into the future, and related to an historic problem around data access that stretches back to the earliest databases.
I’ve written about data access before. I founded Starburst to solve the problem. I care about it deeply, and it informs most of what I do. In fact, in some ways, I’ve dedicated my life to solving the data access challenge and to overcoming the challenges of data centralization, preserving optionality, and encouraging openness.
It was perfect timing then when the AI revolution swept the globe and brought data access needs to the forefront. AI and data access are twin problems. As I’ve written before, AI is only as good as the data it can access. This is true of all data systems, but it’s especially true of AI systems because we ask them to do so much for us, to think like us, to mirror us, to anticipate us, in a way, to be us. And to be us, they need to know us. Knowing us, knowing our organizations and our businesses requires access to the sum total of your organizational data estate. And that requires data access. Not just any data access either, but universal data access.
So we return to the problem. How do you access all of your organization’s data? Two years ago, I wrote the Icehouse Manifesto, outlining why Iceberg + Trino represent the best possible data stack. But more fundamental even than that is data access, so it’s time for a second article, this time a blueprint about data access, the problem of our times. It’s written for any organization approaching this brave new world of AI and encountering the one, universal problem that impacts all AI in production.
How do you access all of your data?
A brief history of the data access challenge
The first thing to know about data access is that this isn’t a new problem. AI might have pushed it to the forefront of our consideration, but the problem is as old as data itself, the problem of centralization and decentralization.
Data centralization and its discontents
From the earliest databases and data warehouses, centralization has been a problem. Data is scattered across different sources, and that presents a critical access problem.
To solve this problem, you basically have two options.
Option 1 – Centralize all your data
The first option is the one that traditionally is pushed as the perceived wisdom within the industry. Take all of your data, wherever it is, and move it, copy it, or otherwise centralize it into one giant data source. This is the approach pushed by the original heavyweights of the industry, Oracle, Teradata, and by new entrants in the market like Snowflake.
The problem with data centralization is that it creates a headache that often cannot be fully solved. Moving data, copying it, is an unending problem for data engineers. Moving it requires a lot of work, effort, and cost.
And the job also isn’t done once you move it. As you centralize, as you get closer to your goal, you incur higher and higher costs. At the same time, you also suffer vendor lock-in and a lack of optionality. This problem only gets worse the further into the tunnel you move. This has been the traditional headache of data engineering for decades, and it hasn’t gone away. AI has only made the problem more apparent and more critical.
Option 2 – Access all of your data using data federation without moving it
The second option is the opposite of the first. Instead of moving all of your data into a central data source, you leave the data where it is and access it from a single location. This sounds good, and it is.
So why haven’t we all taken this approach?
Option 2 requires technology capable of pulling it off. In data architectural terms, it requires data federation. Historically, the performance of data federation was not competitive with data centralization. It required innovations in federation that allow it to perform well enough that federation is competitive.
That’s a technological challenge, but it’s one that we have now solved. In the last few years, we reached that point.
The data federation era has now arrived
The federation era is now here. That simple change, changes everything.
In fact, realizing that we had reached that point is the very reason that I founded Starburst. I saw with my own eyes that data federation and universal access were a total game changer that would upend the entire orthodoxy of data centralization, overturning decades of compromises and stalled projects. Bringing data federation to market in a way that everyone can choose Option 2 has been our singular goal from day one, and that mission has only become more important with time.
Why AI requires data federation to reach its potential
We are now in a position where both of these trends have converged. Data federation technology, like Starburst, has reached a point where it is a competitive and compelling alternative to data centralization. At the very same time, the AI revolution has made data access the precondition for enterprise AI production success. In this, AI has acted as a kind of forcing function for data federation, bringing forward a change that was already in play.
Context and the need for data access
At this point, it’s worth taking a step back and asking why AI needs so much data in the first place?
AI is an amazing revolution that continues to expand and propagate in new and unexpected ways each day. It is, without a doubt, the most important thing to happen to data in a very long time.
But it also has an achilles heel, especially in an enterprise setting. It needs context.
Without context, AI relies on its basic training. That is sufficient for some use cases, but not all. To succeed in the way that enterprise clients expect, the basic training of the LLM model needs to be augmented with context.
What does this context look like?
To answer that, you need to think about what context really is. Context is data, or rather, it’s the data that underlies the assumptions of your business. That seems simple enough, but there’s nothing simple about context.
The problem of context
If I were to ask you, where does your business’ context reside, what would you say? You might point to a production database, or maybe a data warehouse, and that would certainly be some of the context that underpins your business.
But is that all of it?
The thing about context is that not all context is found in one singular place. It’s scattered to the wind, across a dozen source systems, SaaS platforms, and heterogeneous data sources. Often, it’s that way for a reason. Whether for organizational expediency, regulatory compliance and data sovereignty, or legal reasons, contextual data is often held in multiple data sources for good reasons.
Beyond that, a lot of contextual data is not even written down in the first place. It might exist in the form of unwritten, institutional knowledge, or practices known by a team or even an individual. This kind of context isn’t even data in the first place–at least not yet–but it needs to be to give AI what it needs to succeed, particularly in the case of agentic AI.
How the problem of context is really a problem of data access
Seen this way, the problem of context is really a problem of the heterogeneity of data sources coupled with the lack of specific, human knowledge being written down in the first place.
How do you solve for that?
Many organizations are attempting to solve it using Option 1 – the data centralization paradigm. But this doesn’t work. It doesn’t work for the same reasons that it didn’t work in the past, plus a few more.
In the past, centralization failed because it was too labor intensive for its own good. It failed because some data can’t be moved. It failed because it creates a situation of infinite regress, where data centralization becomes a black hole of productivity. That’s bad for analytics, but it’s fatal to AI projects that require context to succeed.
Now consider Option 2 – data federation. It solves all of the issues that AI needs. It allows you to access your organization’s entire contextual data, wherever it lives, and begin using it quickly. It allows you to remain compliant with regulations that inhibit data migration. It allows you to side step the problems of data centralization entirely.
The data centralization era is dead, long live data federation
All of this signals something significant. AI is changing our world, but to do so, it’s running up against the remnants of the data centralization paradigm in a way that is unsustainable.
Something has to give, and that something is data centralization.
In contrast, data federation fits the AI era’s needs exactly. It provides performant, immediate access to data wherever it lives. It provides a way of reaching contextual data that other approaches can’t reach. It offers a universal path to context.
Starburst as the context layer for AI
We’re there already. A universal path to accessing all of your data, all of your context, is exactly what Starburst provides. In this, it is built exactly for the moment that AI demands, a platform capable of reaching everything.
This is what every organization in the world now needs. It’s what data federation can provide. As AI continues to progress from idea to production, that need will only increase. The race is on for context, wherever it can be found. The moment calls out for data federation technology built for this.
What do you need to achieve this? How do you build a data foundation predicated on context? Let’s look at 4 factors that impact success.
Factor 1 – Data access with data federation
AI can’t use context if it can’t access it. For that reason, access needs to be the first tenant of any data foundation. Because contextual data is everywhere, this means federation is a must.
This is the core mission of Starburst as a company, to provide universal data access to all data sources everywhere.
Factor 2 – Alignment of semantic layer to AI with data products
Access is important, but AI also needs to understand what it’s looking at to be useful. This means aligning AI agents with the semantic layer and business logic that underpins your actual business.
Starburst’s data products are uniquely suited to this need. Data products provide curated, governed access to datasets in ways that map the semantic layer of your business to the contextual needs of your AI. This approach reduces hallucinations by curating the context that AI needs to know your business the way you do.
In effect, it brings the logic of your enterprise to your AI agents.
Factor 3 – Iteration of AI projects over time
Things don’t stand still, and that’s as true of AI projects as anything else. Your organization will continue to create new data sources, new semantic meaning, and new context. Iteration, therefore, is essential to the survival of AI. AI built on the semantic layer of your organization 3 years ago will not serve it well today.
Starburst is designed to access all your data in ways that evolve over time. New data sources can be federated alongside old, and new data products can be added to AI workflows.
Factor 4 – Maintenance of the data lakehouse centric data
Federation allows you to access all of your data sources, but that doesn’t mean that you shouldn’t have a preference. I’ve written before how I think that the Starburst Icehouse architecture built on Iceberg represents the best way forward for most new data. So while old data will continue to exist in many sources, new data should preferentially be located in an Iceberg data lakehouse, and that means treating Iceberg as the centerpiece of your data estate.
But managing Iceberg isn’t automatic. Just like everything, it takes work. Features like Starburst data ingestion, and Icehouse LakeOps are designed to make that maintenance easy as you build out your data foundation.
The turning point of our current era
All of this points to a major shift, a turning point, that will rapidly sweep away the data centralization of the past in favor of a federated model of data access. The catalyst for this is the adoption of enterprise AI itself, and there could be no larger push than that, as we race as a world towards this transition.
Given this, I have never been more optimistic than ever at the importance of things that have interested me for years, universal data access, optionality, interoperability. These things were always important, but now, they are inevitable in a way that cannot be ignored. They will reshape our world along with AI, and do so in a way that remakes data along different lines than before. That shift will be significant, and it is already underway. The only question now is how quickly it remakes the world of data and who is an early adopter of this new paradigm.



