What is the Icehouse

An open lakehouse,

with no lock-in.

Icehouse is the name Starburst gave to a simple idea. Pair an open query engine with an open table format and you get all the warehouse-like capabilities you are used to, with no storage lock-in, no metadata lock-in, and no proprietary SQL dialect to learn.

The Icehouse architecture

Open at every layer, including the catalog.

Starburst extends the original idea with a third open layer: the metastore. Any object storage. Any Iceberg-compatible catalog. One open lakehouse.

Your data

S3, ADLS, or GCS. Files stay in your object storage — no copying data into a proprietary vendor format.

built on

Apache Iceberg

Open table format on top of those files: ACID transactions, schema evolution, time travel, and partition evolution. Iceberg was initially built to run on the Trino engine.

built on

Metastore / catalog

AWS Glue, Unity Catalog, Apache Polaris, Starburst Metastore, or any Iceberg REST catalog — your choice, not ours.

built on

Starburst

Distributed SQL, plus federation to the rest of your estate and the performance tuning a lake needs.

Starburst + Apache Iceberg + Metastore Optionality — a high performance, vendor-neutral lakehouse, open at every layer, all the way down to the catalog.

The Icehouse manifesto · February 2024

Forty years of lock-in, and then it changed.

For four decades, data warehouse vendors kept customers locked into proprietary formats that only their own software could read. Cloud data warehouses promised storage and compute separation, but it was always their storage and their compute. Every byte you loaded in was only ever queryable by that one vendor.

Starburst argued a real alternative had finally arrived: an open query engine paired with an open table format, giving companies a genuine warehouse experience on data they actually own and control. Starburst called the combination the Icehouse.

1

Classic data warehouse

Proprietary storage format, proprietary SQL dialect. Your data only speaks to one vendor.

CLOSED
2

Cloud data warehouse

Storage and compute scale independently. Both are still owned by the same vendor.

STILL LOCKED
3

Icehouse

Open query engine, open table format, standard SQL. Any engine can read the data, and no single vendor owns it.

OPEN

Two ingredients

Neither one alone is a warehouse.

Trino and Iceberg were built independently to solve different problems. Combined, they complete each other and add up to a genuine data warehouse experience running directly on the lake.

Trino: the open query engine

Started at Facebook (as Presto) to query hundreds of petabytes with ANSI SQL. Its creators left in 2019 to run it as an independent open-source project and founded Starburst. Today it is used by many companies at large scale.

Iceberg: the open table format

A metadata layer, originally built at Netflix, that sits over files like Parquet and gives them ACID guarantees — letting many engines safely run INSERT, UPDATE, DELETE, and MERGE against the same lake tables. Iceberg was initially built to run on the Trino engine.

Ingredient 1

Trino

Open query engine

Ingredient 2

Iceberg

Open table format

Together

Icehouse

What makes an Icehouse

What has to be true for something to be an Icehouse.

01

No storage lock-in

You bring your own data. Tables live on the object storage you already use — S3, ADLS, or GCS — not a proprietary format.

02

Open engine, open SQL

Query the data with an open engine that speaks open SQL standards. No proprietary dialect to relearn.

03

Iceberg table format

Every table is stored in the open Iceberg format, so more than one engine can read and write it — multi-engine compatibility, by design.

04

Open metastore

Table metadata lives in a catalog you choose, not one bundled exclusively with a single vendor's compute.

Running an Icehouse

Three things any Icehouse needs to actually operate.

The manifesto is the architecture. Turning it into something a team can run day to day takes real operational capabilities — two of which Starburst delivers as fully managed services.

Data ingestion

Streaming and batch data landing in Iceberg tables, continuously and without hand-built pipelines.

Icehouse Ingest

Iceberg data management

Compaction, snapshot expiration, and retention, so tables stay fast and lean as they grow.

Icehouse LakeOps

Data governance

Access control, data lineage, and auditing across every table in the lake.

Built into Starburst

The Iceberg platform

Starburst is the comprehensive solution for Iceberg tables.

Iceberg was initially built for the Trino engine
Fast adoption of spec updates — Iceberg V3 support
Managed features: table maintenance and ingestion
Starburst

Apache Iceberg

Industry-leading performance
Iceberg Lakehouse optionality — any object storage or metastore catalog
AI on Iceberg

What an Icehouse replaces

Any workload you would point at a data warehouse.

01

Business intelligence

Dashboarding and reporting straight off Iceberg tables.

02

Data transformation

Load and transform pipelines that write back into the same open tables.

03

Data-driven applications

In-app analytics served from the lake, not a separate serving layer.

04

AI and ML

Fast, governed access to training and scoring data without a copy step.

Further reading

Go deeper on Icehouse.

The Icehouse Manifesto: Building an Open Lakehouse

Read the post

Automating the Icehouse: Fully Managed Open Lakehouse

Read the post

What is an Icehouse: Trino, Iceberg, Data Lakehouse

Read the post

Next Gen Data Management with Icehouse Architecture

Read the post

CDC with Trino and Iceberg

Read the post

Building a lakehouse with dbt and Trino

Read the post

Go deeper on the two managed capabilities

Icehouse Ingest

Stream Kafka topics and land S3 files into query-ready Iceberg tables.

Explore Icehouse Ingest

Icehouse LakeOps

Serverless table maintenance and lake observability for every Iceberg table.

Explore Icehouse LakeOps

Open architecture. Open standards. Your data.

Start building your Icehouse on Starburst Galaxy today.