Why Enterprise AI Success Comes Down to Data Access

And how Starburst is solving the problem

Share

Linkedin iconFacebook iconTwitter icon

More deployment options

AI agents (still) have a production problem. In 2025, it was estimated that some 88% of agent projects never reached production. It’s 2026, and the most positive estimates put deployment at just over 30%.

The wave of stalled AI agent projects may seem surprising. But if you’ve worked in the data world long enough, it feels familiar.

We’ve seen this before, just in different forms, most notably around data centralization projects. Think of the multi-year data warehouse migration, the one data lake to rule them all, or the cloud replatforming project. Each new wave promised a solution. Each one fell short due to the monolithic needs of a one-size-fits-all solution.

AI is still new. It’s improving and changing rapidly. But the need to avoid monolithic mindsets remains constant. How do you do that? Well, all of the monolithic projects have something in common. They all approach data access in the same way, through a need to centralize everything first, before you can get started. This is the same with AI, and the solution to doing AI the right way also starts with data access. 

In this sense, the problem was never the tooling or the model. AI is hitting the same data access wall that analytics and data apps projects have always faced. It’s just hitting it faster and at greater cost. In this article, we’ll dig into this persistent problem and how to solve it once and for all.

Why the centralization mindset is voracious and unendingly problematic 

Let’s look at centralization as a problem in itself. It’s nothing new. AI agents aren’t the first workloads to need massive amounts of data. Analytics and data apps have long needed to reconcile and transform disparate data from across the enterprise.

For 20 years, the solution to growing data demand has been the same. Move everything into a single system to make working with it fast and easy.

That thing morphed over the years, from Hadoop to the cloud data warehouse and then the data lake and lakehouse. Each of these solutions solved a real problem, such as consolidating distributed data and making data processing at scale cost-efficient.

But none of them ever got us to the promised destination,  a “single source of truth” for everything.

Meanwhile, data volumes and sources exploded. The mean number of data sources at an enterprise was 400 in 2020. Some companies had over 1,000. Most are likely managing far more than that now. The explosion sent analysts and data engineers on a mad hunt for data from new SaaS tools, acquired companies’ systems, regional systems, and data that couldn’t legally move.

Unfortunately, we were all still stuck in the centralization mindset. 

AI creates the same data access problem, but multiplies it

AI is different in so many ways. AI agents are powered by large language models (LLMs), neural networks trained on large volumes of general data. They’re great at parsing human-language requests and generating outputs. But they know little to nothing about your business.

Teaching AI agents about your business requires giving them access to context. What is context? Well it’s data, really. Specifically it’s any data that provides critical information about your business.

Seen this way, the problem of accessing context is really just data access by another name, because context can live anywhere in your enterprise. That’s one of the problems with it. It might be your SaaS marketing tools, a data warehouse, emails and PDFs, an internal wiki, an Amazon S3 bucket, and any of few dozen different data sources.

Missing context, in other words, is simply data that your AI can’t reach yet. This problem was already bad for analytics and data apps, but AI agents make it 10 times worse because enterprises are increasingly relying on them to make autonomous decisions on users’ behalf.

This is why so many AI agents never make it to production scale. Gartner estimates that a lack of AI-ready data threatens to put 60% of projects at risk.

What impacts contextual data access? 

There are three barriers that prevent data access, including: 

Fragmentation. The data that carries business meaning is scattered across legacy on-premises systems, multiple clouds, and repositories that many in the company don’t even know exist.

Format. Much of this data lives in unstructured and semi-structured formats that data pipelines weren’t built for: PDFs, chat logs, emails, transcripts, and video. Most of this data couldn’t even be accurately parsed before the advent of AI. This data must be mined and made consistently available to AI agents.

Governance. From a security standpoint, most data repositories aren’t set up to provide autonomous AI agents with well-scoped access to properly sanitized data.

The autonomy of AI agents makes data access a table-stakes problem. A human can compensate for a failed data dashboard that looks “off.” An AI agent can make million-dollar mistakes in seconds.

Why you can’t centralize your way out

The reflexive solution to all of this? Centralization, of course. This time, with AI. But centralization hasn’t worked before. It won’t magically start working now.

Centralization was never practical, not for data warehouses, not for data lakes, not for the cloud, and not for AI. As before, migrating everything isn’t just expensive; it’s impossible. Your business is creating exponentially more data every year. 90% of the data that exists now is only two years old. A “migration” would be never-ending.

Migrations are measured in years, and many literally never finish. Meanwhile, AI projects have to launch by end of quarter. The world’s moving too fast and producing too much data to wait on creating a “single source of truth.”

Why an enterprise intelligence platform makes your company the source of truth

The solution isn’t to try to create yet another single source of truth. It’s to make your company itself the source of truth. Instead, what you need is a solution that moves at the speed of your business. 

This requires one key ingredient, universal data access.

In a universal data access model, you don’t centralize everything up front. We know that doesn’t work. Instead, you empower users to discover and access data where it lives today.

Traditional data pipelines are a process-based solution to the data access problem. Universal data access is a technology-based solution. It turns data access from a series of tasks managed in a backlog, but into an architectural capability available on demand.

This doesn’t mean you never centralize. It means you can selectively centralize. If a given dataset requires it, you can move it into an efficient, modern data format, such as Apache Iceberg. That’ll give you all the benefits of a format like Iceberg, such as improved performance, advanced features (like time travel and rollback), and stronger governance capabilities.

Building out universal data access doesn’t require creating a new data architecture from scratch. It requires an enterprise intelligence platform like Starburst. An enterprise intelligence platform provides a set of services that enable managing data for AI workloads at scale:

Federated access to every source. Finding and querying data across the enterprise requires not a single source of truth, but a single access point. Data federation enables users to pull data using distributed queries, no matter where it lives. This brings data to AI workloads now, not months from now.

Governance that’s built in. Data access can’t scale unless data is well-governed. This requires fine-grained data access controls, data lineage to validate the source and provenance of data, and centralized audit logging applied at the access layer itself.

Meaning attached to the data. Context is more than just raw data. It requires meaning. This takes the form of an enterprise context layer that sits between your agents and your data, and a semantic layer that standardizes the business’s key metrics.

Starburst is built to solve data access at the root of your data foundation

The context problem isn’t isolated to a single company. It plagues the entire AI industry, and that means that the entire world has a data access problem.

The problem is the one that Starburst was built to solve. The Starburst Enterprise Intelligence Platform brings your data to AI by using:

  • An Analytics Engine that powers federated access across on-premises, cloud, and hybrid environments, with no data replication required
  • An Enterprise Context Layer providing the definitions, policies, and metadata that turn raw sources into curated, trustworthy Data Products for AI
  • An Agentic Control Plane that provides governance and guardrails to coordinate AI workflows, so every agent query is permitted, logged, and auditable

Image depicting the 4 different layers of the starburst enterprise intelligence platform.

For users, the Enterprise Intelligence Platform also supports the Starburst AI Data Assistant (AIDA), a conversational, natural-language interface to your organization’s distributed data. Meanwhile, our hosted Model Context Protocol (MCP) server ensures that AI agents can access governed data autonomously via an open standard.

The winners of the AI era will be the companies whose agents can reach, trust, and act on their data, no matter where it lives. Discover how Starburst can get you there by asking us for a demo today.

Start for Free with Starburst Galaxy

Try our free trial today and see how you can improve your data performance.
Start Free