
Data federation is changing in response to AI. What was always a useful way of accessing multiple data sources is now the default approach enterprises reach for when they need governed, proving the pathway for universal, AI-ready data access to data spread across dozens of systems. That includes data lakes and SaaS applications, without duplicating the data first. As such, federation has become the standard layer for querying everything, replacing a good idea for analytics with an essential for AI.
This article covers what changed technically between 2020 and 2026, why AI agents made federation urgent rather than optional, and how regulations like DORA make a new case for analytics and AI using federation. It also covers how federation fits alongside the data lakehouse rather than replacing it, plus a practical list of what to evaluate in a federation platform today and where the approach still runs into limits.
Key takeaways
- Data federation now functions as production infrastructure for AI agents.
- Data sovereignty regulation, including GDPR’s cross-border transfer rules and DORA’s third-party vendor location requirements, pushes enterprises toward analytics in place instead of full centralization, turning federation into a compliance requirement rather than an optional convenience.
- Federation and the data lakehouse work together. The data lakehouse handles storage and governance for the data you centralize, and federation extends that reach across everything else.
- The open question for most buyers is where federation still adds more latency or complexity than a consolidated store would.
What data federation means now, versus five years ago
Five years ago, data federation mostly meant data virtualization, a thin query layer that let analysts avoid copying every dataset into a data warehouse before running a report.
Today, data federation gives you a unified, virtual view of data across disparate sources, and it’s built to handle production workloads rather than ad hoc reporting. The core idea hasn’t changed. You still query across all those data sources as if they were a single database rather than copying data into a central store before you can query it. What’s different is that the engines underneath now handle enterprise scale, governance, AI workloads, and business intelligence queries all at once.
In recent years, the average data estate has only grown more fragmented. Enterprises now run operational data across cloud platforms, on-premises systems, and dozens of SaaS applications, and no single team can justify centralizing all of it before every question gets answered. Federation gives you a way to answer those questions without that up-front migration project.
Why AI agents made 2026 the tipping point for federation
AI agents have helped to make made data federation the default. Agents need current data instead of a snapshot from last night’s batch job. They also need it from every system that holds relevant context, including CRM records, ticketing systems, and operational databases alongside the data warehouse.
That requirement exposed a hard limit in the data centralization model. Copying every relevant datasource into one data warehouse before an agent can use it adds latency, cost, and a growing list of ETL pipelines to maintain. With federation, an agent can query a datasource directly, through a governed context layer that AI agents and BI tools share, without waiting for a new pipeline to copy that data somewhere else first. Production workloads still benefit from caching or materialized views for speed, but that happens inside the platform rather than as a separate project someone has to build.
That governed access increasingly runs through the Model Context Protocol (MCP), the mechanism most agent frameworks now use to reach data and tools directly instead of a custom integration for every system. MCP moved to the vendor-neutral Agentic AI Foundation under the Linux Foundation in December 2025, and Anthropic, OpenAI, Google, and Microsoft all support it natively, making it the closest thing agent-to-data access has to a default standard in 2026.
Why most agent projects still don’t reach production
The importance of data access is already showing up in deployment numbers. Deloitte’s 2026 State of AI in the Enterprise research found that only 25 percent of organizations moved 40 percent or more of their agent pilots into production. On a different measure, S&P Global Market Intelligence and McKinsey put the share of enterprises with at least one agent already in production at 31 percent in 2026. Data access is a recurring reason why most agent projects never reach production when the underlying data is scattered, ungoverned, or too slow to reach. With federation, agents get live access without a new copy pipeline for every datasource.
From data virtualization’s reputation problem to production-ready federation
Data virtualization earned a reputation for being slow and brittle, and that reputation followed federation for years even as the underlying technology improved. Early implementations struggled to push complex operations down to the datasource, so queries often pulled far more data than necessary before filtering it, which made performance unpredictable at scale.
The technical changes behind production-ready federation
Three technical changes closed those limitations.
- Query pushdown got substantially better, so filtering and aggregation now happen at the datasource instead of after every row crosses the network.
- Connector depth expanded. For example, Starburst’s connector ecosystem now spans 50 or more data sources across cloud platforms and on-premises systems, and it supports major table formats, including Iceberg and Delta Lake, on Amazon S3, Azure Blob, and Google Cloud Storage.
- Platforms like Starburst now allow for a semantic and context layer to sit on top of the query engine, giving both AI agents and BI tools a consistent, governed view of what each field means and where it lives.
Governance and data sovereignty as a forcing function
One data warehouse used to mean one set of access controls and one audit trail. In 2026, consistent governance increasingly comes from keeping data where it is rather than from centralizing it.
Data sovereignty regulation pushes data federation over centralization
Data-residency and sovereignty rules are a growing reason enterprises keep analytics in place instead of centralizing everywhere. GDPR’s cross-border transfer restrictions and national data-localization rules increasingly require financial and public-sector data to stay inside a jurisdiction. The Digital Operational Resilience Act (DORA), which took effect in January 2025, adds a narrower but related pressure on EU financial institutions. It requires them to maintain a register of where third-party ICT vendors process and store their data and to disclose location changes in advance. Together, these rules make it harder to justify copying transaction data into a single region for analysis. Instead, institutions increasingly run analytics in place, without moving data across borders, which keeps sensitive records under local control while still allowing enterprise-wide queries. Regulatory pressure like this is pushing federation from a nice-to-have into a compliance requirement for global organizations.
Data federation and the data lakehouse work together
Importantly, federation hasn’t replaced the data lakehouse, and that was never the goal. The data lakehouse still supplies durable storage, transactional guarantees, and governance for the data you choose to centralize, while federation supplies reach across everything else. That division of labor is a starting point rather than a hard line. Federation engines, including Starburst, increasingly materialize results into lakehouse table formats like Iceberg for performance, and lakehouse platforms are adding native federation features of their own. Even so, federation is what covers eliminating data silos without another copy of the data in systems that will never move into the data lakehouse, such as operational databases and SaaS applications.
In practice, most enterprises run both. You centralize the data that benefits from a single, well-governed source of truth, and you federate the rest, since a data lakehouse and federation solve different problems within the same data estate rather than compete as alternative approaches.
What to evaluate in a federation platform today
A handful of criteria separate production-ready federation from tools that are not up to the job. First, check the number of connector connectors and table format support, since a platform that only reaches your data warehouse isn’t federating anything. Second look at query performance, which often depends on how much work the platform pushes down to the datasource rather than pulling raw data across the network. Governance features need the same scrutiny. Look for row-level and column-level access control that’s consistent across every connected datasource.
Finally, ask whether AI agents get the same governed access that BI tools already have, and ask specifically whether the platform supports the Model Context Protocol (MCP), the standard most agent frameworks now use to connect to governed data and tools. If querying everything without copying it between systems only works for dashboards, not for the agents your teams are building, you’re evaluating yesterday’s federation platform. Vendor claims about connector counts are easy to make. Ask for the specific list of data sources and table formats that matter to your stack, rather than taking a marketing number at face value.
Next steps
Want to know more about how federation fits into a modern data setup? Check out this webinar to see how query engines, data lakehouses, and AI fit together.
FAQs
What is data federation?
Data federation gives you a way to query data across multiple systems as if they were one database, without copying everything into a central store beforehand. It works through connectors that reach into each datasource, including data lakes, data warehouses, and SaaS applications. A query engine then pushes filtering and aggregation down to where the data already lives.
Why did AI agents make data federation urgent in 2026?
AI agents need current data from every system that holds relevant context, rather than a snapshot from an overnight batch job. Copying each datasource into a data warehouse before an agent can use it adds latency and a growing list of pipelines to maintain, so federation gives agents live, governed access instead.
Does data federation replace the data lakehouse?
No, data federation uses the data lakehouse and the two are not in opposition. The data lakehouse still supplies durable storage, transactional guarantees, and governance for the data you choose to centralize, while data federation provides universal access to all data sources, including the data lakehouses, data warehouses, data lakes, operational databases, and SaaS applications. Most enterprises run both, and current platforms increasingly blend the two, for example by materializing federated results into lakehouse table formats for performance.
What should you look for in a data federation platform?
Start with connector depth and table format support, since a platform that only reaches your data warehouse isn’t federating anything. Additionally, check how much filtering and aggregation the platform pushes down to the datasource. Confirm that AI agents get the same governed access to data that BI tools already have.



