5 Questions Every Dremio Customer Should Ask About the SAP Deal

A practical checklist for protecting your data strategy while the acquisition settles

Share

Linkedin iconFacebook iconTwitter icon

More deployment options

The SAP acquisition changes something important for Dremio customers, though it is probably not the thing you would expect.

First, let’s look at what doesn’t change. The acquisition doesn’t change your clusters. They will run today exactly as they did last week, and the deal does not force you to do anything right now. For many teams, continuing to run what already works is the sensible near-term call.

Thinking long term about Dremio and SAP as a business choice

What does change though is both quieter and more long term. Imagine the next year or two of ordinary decisions. A new question needs a source you have not connected yet. A new AI project needs data that lives outside the lake. Each time, there is a path of least resistance, and it keeps pointing in the same direction, toward the platform’s own world and away from everything else. That gradual drift is the thing worth watching, and it is the reason the questions in front of you have quietly changed.

Does data centralization still make sense today? 

Ultimately, as in any acquisition, the platform you originally chose as an independent product now sits inside one of the largest software portfolios in the world, and large portfolios exert a consistent pull. This business model is based around data centralization and is designed to reward bringing more of your data into their ecosystem, and they make everything outside that ecosystem a little harder to reach.

That focus is not new, and it is not really about SAP. It reflects a much older habit, the assumption that data has to be moved into one place before it can be useful. That assumption made sense in an era when querying across systems was genuinely difficult, and for years it was simply how things were done. But it carries costs that are easy to stop noticing over time, in the copies you maintain, the pipelines you build, and the data that is always a step behind its source, and those costs accumulate faster than most teams expect. An acquisition is a good moment to weigh them honestly, because the gravity of a larger platform is precisely what turns  the drive to query  into a push towards centralization.  

How to approach the Dremio question systematically

So this is not a case for ripping anything out, and it is certainly not a case against SAP. Instead, it is a checklist. There are five questions, and they build on one another, from the roadmap you were promised through to whether your AI can be trusted at all. The time to work through them is now, while the deal is still settling and you still have room to choose.

Question 1: What happens to the roadmap you were promised?

Today, Dremio sets its own priorities. Once the deal closes, those priorities have to compete for budget and attention inside a much larger organization, one with its own products to sell and its own reasons to encourage consolidation.

This is not a claim about bad intent. It is simply how acquisitions reshape a roadmap, and it means the honest question is a practical one. Which of the features you were counting on are still funded, and on what timeline?

There is one dynamic worth watching in particular. When the capabilities you care about serve the broad market, they tend to get built into the product. When they serve the parent company’s consolidation story instead, they tend to get built faster. Those two are not always the same, and when they diverge, the engineering usually follows the parent.

That is the abstract version of the risk. It becomes concrete the moment you look at where your data actually lives.

Question 2: How much of your data lives outside SAP, and can the platform still reach it?

This is the foundation the other questions rest on, and it is easily underrated if most of your analytics currently sit inside the lake.

Start by mapping your data sources honestly. Count all the databases, data warehouses, object stores, and SaaS systems that live outside any SAP product, and then ask what querying them across systems is likely to look like over time, on a platform whose center of gravity now sits inside a larger ecosystem.

Again, this is not about bad intent. It is about where the easy path leads. The in-ecosystem path keeps getting smoother while everything else receives a little less attention, until eventually the easy option and the in-ecosystem option are the same thing. A platform that genuinely reaches your whole estate is one that keeps any single vendor, SAP included, as one source among many rather than the gravitational center.

There is a blunter version of this question that practitioners will recognize immediately. Every source the platform cannot reach natively is one more pipeline someone has to build and maintain. So it is worth asking where the connector investment is actually going, before the drive to centralize becomes the default answer to every new request.

Reach tells you what you can get to today. The next question is whether you can still leave.

Question 3: Are you tied to a single table format, and what would leaving cost?

Open table formats were supposed to protect you from lock-in. But a platform built around a single format only delivers that protection inside that format’s own ecosystem.

So if your data is committed to one format under a single vendor’s catalog, the question worth asking directly is what it would actually cost to move a meaningful workload elsewhere, and whether that cost is rising over time.

Portability was always the point of going open, and portability is measured on the day you try to leave, not the day you sign a new contract. A useful test is whether your data remains reachable by engines other than the one that wrote it. Open architecture is what preserves that choice, and if the answer to the test is no, or increasingly no, then you have less optionality than the word “open” originally implied.

Ultimately, reach and portability are about getting to your data and keeping it free to move. The next question is what happens once that data is spread across many systems, because that is where control becomes genuinely difficult.

Question 4: Does your governance hold across every source, or only inside the lake?

Governance that works beautifully inside the lake but stops there is not really enterprise governance. It is data lake governance only, and the distinction matters more than it first appears.

As AI agents begin querying across data sources and technologies, that edge becomes the weak point in the whole arrangement. It is where sensitive data can leak, and it is where two agents can reach different conclusions from the same underlying numbers, simply because different rules applied on different sources.

The question to ask is whether your access controls, business definitions, and audit trails travel with the data across every source you query, or whether they only hold within the platform’s own storage. Consistent governance across the whole estate is what allows you to trust an answer regardless of which systems it touched along the way. It is also precisely the kind of cross-source capability that a consolidation-focused roadmap has little incentive to prioritize.

Everything to this point has been about your data. The final question concerns what happens when something begins reasoning over that data on its own.

Question 5: Can your AI workloads reach all of your data, or just the part nearest SAP?

AI is where narrow reach stops being an inconvenience and becomes a larger problem.

An agent reasons over whatever it is able to query, and it reports its answer with exactly the same confidence whether it saw all of your data or only a slice of it. A system it cannot reach is not a system it flags for you. It is one it ignores silently.

So the real question for any platform carrying your AI workloads is this. Can an agent reach the databases, warehouses, and SaaS systems it may need in the middle of its reasoning, at the concurrency a fleet of agents generates, and against current data rather than a scheduled copy?

If the reachable set turns out to be “your data, as long as it already lives in the right place,” then your agents will produce confident answers built on a partial picture, and neither you nor they will necessarily notice. Consider an agent assessing churn risk that can see the CRM but not the support tickets. It does not hand you a partial answer wrapped in a caveat. Instead, it hands you a complete-sounding one, delivered in the same confident tone it uses when it is right. That is a far worse failure than a slow query, because a slow query at least tells you that something is wrong.

This is also the point where moving data first stops being merely slow and becomes actually unworkable. An agent asks a question that did not exist five seconds ago. Waiting to move, copy, and re-govern the relevant dataset before it can answer is not a realistic option at that speed. Either the data reaches the agent now, wherever it happens to live, or the agent answers without it.

You do not have to centralize data to get value from your platform

Follow those five questions to their conclusion and they all point in the same direction. A roadmap you no longer control leads to reach you slowly lose, which leads in turn to data you cannot leave with, governance that stops at the lake’s edge, and finally an AI you cannot fully trust. That is the full cost of the drift, if you allow it to run unchecked.

Avoiding that outcome is not a binary choice. It was never a matter of staying all-in on Dremio under SAP or replacing everything and starting over. 

There is a third path, and every one of these questions points toward it. Keep what works, and add reach around it.

If Dremio is serving a genuine lakehouse workload well, there is no urgency to remove it, and if the acquisition brings tighter integration for your SAP data, that is a real gain worth keeping. None of this requires a position as extreme as never moving data. Moving data is sometimes exactly the right call. The point is narrower and more useful than that, which is that data movement should be a deliberate decision rather than the automatic first step. Move the data that genuinely needs to move, and stop moving the rest by default.

How to handle the data that lives outside your data lake is the key question, especially due to the AI workloads that need to span both worlds. A data federation layer sits alongside your existing environment rather than replacing it, so your SAP and Dremio investments keep doing their job while your analysts and agents gain governed access to the rest of the estate while it stays where it is.

For the teams maintaining all of this, that means less pipeline and more answers. For the architects, it means the freedom to use the platforms you already run without funneling every workload into a single destination first. Open architecture is what protects that choice.

That is the posture worth adopting while the deal settles. Keep your options open, protect your access to the data that lives outside any one vendor’s walls, and make sure your strategy remains a decision rather than a default.

If you want to pressure-test where your current setup leaves you, the full Starburst and Dremio comparison is a reasonable place to start. And the questions above are worth putting to any platform you depend on, this one included.

Want to know more? Sign up for the webinar where we discuss Dremio, SAP, and Starburst

 

Start for Free with Starburst Galaxy

Try our free trial today and see how you can improve your data performance.
Start Free