An open lakehouse,
with no lock-in.
Icehouse is the name Starburst gave to a simple idea. Pair an open query engine with an open table format and you get all the warehouse-like capabilities you are used to, with no storage lock-in, no metadata lock-in, and no proprietary SQL dialect to learn.
The Icehouse architecture
Open at every layer, including the catalog.
Starburst extends the original idea with a third open layer: the metastore. Any object storage. Any Iceberg-compatible catalog. One open lakehouse.
Your data
S3, ADLS, or GCS. Files stay in your object storage — no copying data into a proprietary vendor format.
Apache Iceberg
Open table format on top of those files: ACID transactions, schema evolution, time travel, and partition evolution. Iceberg was initially built to run on the Trino engine.
Metastore / catalog
AWS Glue, Unity Catalog, Apache Polaris, Starburst Metastore, or any Iceberg REST catalog — your choice, not ours.
Starburst
Distributed SQL, plus federation to the rest of your estate and the performance tuning a lake needs.
Starburst + Apache Iceberg + Metastore Optionality — a high performance, vendor-neutral lakehouse, open at every layer, all the way down to the catalog.
The Icehouse manifesto · February 2024
Forty years of lock-in, and then it changed.
For four decades, data warehouse vendors kept customers locked into proprietary formats that only their own software could read. Cloud data warehouses promised storage and compute separation, but it was always their storage and their compute. Every byte you loaded in was only ever queryable by that one vendor.
Starburst argued a real alternative had finally arrived: an open query engine paired with an open table format, giving companies a genuine warehouse experience on data they actually own and control. Starburst called the combination the Icehouse.
Classic data warehouse
Proprietary storage format, proprietary SQL dialect. Your data only speaks to one vendor.
Cloud data warehouse
Storage and compute scale independently. Both are still owned by the same vendor.
Icehouse
Open query engine, open table format, standard SQL. Any engine can read the data, and no single vendor owns it.
Two ingredients
Neither one alone is a warehouse.
Trino and Iceberg were built independently to solve different problems. Combined, they complete each other and add up to a genuine data warehouse experience running directly on the lake.
Trino: the open query engine
Started at Facebook (as Presto) to query hundreds of petabytes with ANSI SQL. Its creators left in 2019 to run it as an independent open-source project and founded Starburst. Today it is used by many companies at large scale.
Iceberg: the open table format
A metadata layer, originally built at Netflix, that sits over files like Parquet and gives them ACID guarantees — letting many engines safely run INSERT, UPDATE, DELETE, and MERGE against the same lake tables. Iceberg was initially built to run on the Trino engine.
Ingredient 1
Trino
Open query engine
Ingredient 2
Iceberg
Open table format
Together
Icehouse
What makes an Icehouse
What has to be true for something to be an Icehouse.
01
No storage lock-in
You bring your own data. Tables live on the object storage you already use — S3, ADLS, or GCS — not a proprietary format.
02
Open engine, open SQL
Query the data with an open engine that speaks open SQL standards. No proprietary dialect to relearn.
03
Iceberg table format
Every table is stored in the open Iceberg format, so more than one engine can read and write it — multi-engine compatibility, by design.
04
Open metastore
Table metadata lives in a catalog you choose, not one bundled exclusively with a single vendor's compute.
Running an Icehouse
Three things any Icehouse needs to actually operate.
The manifesto is the architecture. Turning it into something a team can run day to day takes real operational capabilities — two of which Starburst delivers as fully managed services.
Data ingestion
Streaming and batch data landing in Iceberg tables, continuously and without hand-built pipelines.
Icehouse IngestIceberg data management
Compaction, snapshot expiration, and retention, so tables stay fast and lean as they grow.
Icehouse LakeOpsData governance
Access control, data lineage, and auditing across every table in the lake.
The Iceberg platform
Starburst is the comprehensive solution for Iceberg tables.
Apache Iceberg
What an Icehouse replaces
Any workload you would point at a data warehouse.
01
Business intelligence
Dashboarding and reporting straight off Iceberg tables.
02
Data transformation
Load and transform pipelines that write back into the same open tables.
03
Data-driven applications
In-app analytics served from the lake, not a separate serving layer.
04
AI and ML
Fast, governed access to training and scoring data without a copy step.
Further reading
Go deeper on Icehouse.
The Icehouse Manifesto: Building an Open Lakehouse
Read the postAutomating the Icehouse: Fully Managed Open Lakehouse
Read the postWhat is an Icehouse: Trino, Iceberg, Data Lakehouse
Read the postNext Gen Data Management with Icehouse Architecture
Read the postCDC with Trino and Iceberg
Read the postBuilding a lakehouse with dbt and Trino
Read the postGo deeper on the two managed capabilities
Icehouse Ingest
Stream Kafka topics and land S3 files into query-ready Iceberg tables.
Explore Icehouse IngestIcehouse LakeOps
Serverless table maintenance and lake observability for every Iceberg table.
Explore Icehouse LakeOpsOpen architecture. Open standards. Your data.
Start building your Icehouse on Starburst Galaxy today.