
The conversation around enterprise AI has fundamentally changed. We have moved past the phase of experimentation and into the harder, more consequential phase of engineering the AI leap. What does enterprise agentic AI needs to succeed in production? Ultimately, it comes down to three things:
- Accuracy
- Consistency
- Auditability
An AI agent does not hallucinate because the model is weak. It hallucinates because the data beneath it is stale, the definitions are inconsistent, or the path it took cannot be traced. As production agentic AI, real-time log analytics, and high-concurrency analytical workloads scale, data platform teams need more than just a basic query engine. They need unified lakehouse management, real-time operational observability, zero-downtime engine stability, near real time log/event processing, and modern developer tooling.
With this release, Starburst Galaxy delivers the engineered data foundation required for production AI and analytics across six core pillars.
Agentic trust, intelligence, and federated business context
Enterprise AI and automated data pipelines require data that is accurate, governed, and rich in business context.
Galaxy delivers this through Starburst AIDA, which now includes native AI Safety Guardrails, programmatic Data Quality APIs, and Editable Metadata Management.
Federated business context and metadata management
Editable Metadata Management and Public Context APIs
This feature enables data engineers and platform teams to read and write business descriptions, custom metadata, and table-level business rules directly in the Catalog Explorer or programmatically via public REST APIs. Bidirectional sync between Catalog Explorer assets and Data Products eliminates metadata silos, while dbt integrations allow descriptions, tags, and column mappings to flow automatically into Starburst to ground AI reasoning.
Rich Semantic Guardrails for AI
Semantic Guardrails expands the enterprise context layer with description fields specifically designed to improve context capturing. Curators can define exact join key pairs, join cardinalities (1:1, 1:n, n:n), Fact vs. Dimension tags, valid value allowlists, column synonyms, and suggested questions directly on datasets to help reduce agent hallucination during complex joins and aggregations by providing the necessary ground truth.
Programmatic Data Quality APIs
Programmatic data quality APIs deliver full read and write API support for Starburst Galaxy Data Quality checks. Platform teams can programmatically create, manage, and monitor data quality tests as code in GitHub, as well as link tests directly into Airflow orchestration DAGs to halt downstream jobs automatically if quality thresholds fail.
AIDA: Trusted conversational analytics on governed data
While AI agents promise to put analytics in the hands of every business user, they are only as trustworthy as the data and controls behind them. This release expands Starburst’s AIDA with additional ways to extend the agent through AIDA Studio and new audit and usage visibility for administrators, all built on governed data products and the access controls already in place in Galaxy.
Analytics Agent Experience
Natural language analytics using data products
Lets business users ask questions of curated data products in natural language, without writing SQL, and follow up to dig into what is driving a result. Every answer shows the SQL and steps behind it, so users can see exactly how a result was produced.
Role-Based Access Enforcement
Runs every AIDA query under the role the user is signed in with, so the same access controls that govern data everywhere else in Galaxy apply to every answer AIDA returns.
Personas, chat history, and interactive charts
Built-in personas (Executive, Analyst, and Data Engineer) and custom personas tailor how AIDA responds to each audience. Chat history lets users pick up an investigation where they left off, and interactive charts turn query results into visuals.
AIDA Studio
Skills
Packages the instructions for recurring analytical work, so everyone who runs a common process follows the same steps.
External MCP Servers
Connects AIDA to tools and context outside Starburst, with per-tool approval required before AIDA invokes any tool.
Governance and Monitoring
AI Guardrails
Default-on protections that keep AIDA from revealing its own instructions, direct it to treat instructions in query results and tool responses as data, keep its built-in tools within the session’s data product, and reject oversized prompts before they reach the model. Optional topic filtering blocks subjects outside AIDA’s intended use. Learn more about AI guardrails.
Agent Audit Logs
An immutable record of every AIDA interaction, including the user, question, data product, persona, steps taken, skills used, and any user feedback. Administrators can filter logs by date, user, data product, and status, and export them as CSV for compliance and reporting.
Usage and Cost Monitoring
Daily AIDA token usage and cost with a per-user breakdown, exportable for internal reporting and budgeting.
The unified Starburst Icehouse console and advanced Iceberg engine scale
While Apache Iceberg has become the open standard for the modern data lakehouse, managing Iceberg in production has historically been fragmented and opaque. This release introduces The Starburst Icehouse Console, a unified, end-to-end Iceberg management experience in Galaxy backed by deep, enterprise-grade Iceberg engine optimizations.
Icehouse Console and Management UX
Dedicated Icehouse Control Plane
A centralized control plane in Galaxy where data teams can discover, optimize, monitor, and govern all Apache Iceberg tables across every catalog without context-switching between tools.
Automated Schema Discovery and Drift Management
Automatically infers and registers Iceberg table schemas during catalog creation, with continuous background schema monitoring to catch and absorb schema drift before downstream BI dashboards or AI models break.
Actionable Iceberg Observability and ROI Metrics
Real-time visibility into table health, file count reductions, storage space reclaimed, and maintenance execution history. Visual cost savings and performance baseline comparisons quantify the exact ROI of your lakehouse upkeep directly to leadership.
Iceberg Engine Scale and Performance
Iceberg Incremental Materialized View Refresh
Introduces scheduled, append-only incremental refreshes for Iceberg materialized views using a user-defined incremental column. By processing only newly added rows across any Trino-queryable source rather than executing full recomputes, Galaxy dramatically shortens refresh windows and eliminates the primary blocker keeping legacy pipelines on Hive.
Engineered Serverless Table Maintenance
Offloads memory-intensive compaction, orphan file removal, and snapshot expiration off standard query clusters onto a dedicated, serverless execution engine. Resumable compactions, progress reporting, and auto-tuning thresholds eliminate small-file fallout after heavy MERGE operations, preventing cluster Out-Of-Memory risks and protecting production query SLAs.
Iceberg Distributed Metadata Retrieval and Planning
Offloads manifest parsing and split generation from the coordinator to worker nodes for queries scanning tables with tens of millions of files. By delegating metadata tasks across cluster workers, Galaxy slashes coordinator memory and CPU consumption, eliminates coordinator bottlenecks on massive datasets, and enables 2x smaller coordinator node sizing for massive cost savings.
Complete Iceberg v3 Spec and Cloud-Native REST Integration
Building on Iceberg v3 deletion vectors and branching, this release expands native v3 type mappings (including Geometry and Geography data types) alongside turnkey Iceberg REST Catalog integrations featuring short-lived credential vending and native cloud auth flows across AWS, Azure, and GCP.
High-Concurrency, data federation, sub-second log analytics, and dynamic warp speed caching
Analytical workloads are expanding beyond traditional BI reporting into massive-scale log, event, and real-time streaming analytics. This is being driven by data federation, capable of accessing data wherever it lives. Engine updates in this release optimize cluster coordination, caching efficiency, and single-table aggregation patterns to deliver maximum price/performance across demanding enterprise environments.
Engine analytics optimizations
Near real time log and event analytics performance
Target engine optimizations focus on high-velocity query patterns, specifically single-table aggregations, heavy filtering, and projections typical of ClickBench benchmarks. By removing internal query bottlenecks, Galaxy enables near real time responses for log and event analytics workloads, directly outperforming dedicated engines like ClickHouse and StarRocks while keeping data safely in open lakehouse storage.
Global dynamic FS caching layer for Warp Speed
Upgrades Warp Speed with a unified, global file system caching layer across multi-catalog deployments. SSD disk space is now dynamically allocated across catalogs in real time based on active query demand, eliminating static per-catalog disk partitioning, maximizing SSD utilization, and accelerating query performance across index-assisted lakehouse catalogs.
Cluster concurrency and routing
Optimized engine coordination for 2-3x higher QPS
Engine refinements significantly reduce cluster-level coordination overhead across worker nodes. A single cluster can now process 2 to 3 times higher query concurrency and throughput, reducing the operational overhead of running dozens of separate clusters for high-concurrency enterprise applications.
Smart load balancing via stateful atomic routing
Replaces random query distribution with intelligent, metric-driven routing across multi-cluster deployment sets. Smart Load Balancing continuously evaluates real-time queue depth and active query counts across clusters. Using a stateful, atomic dispatch counter, it deterministically shuffles concurrent query bursts across healthy deployments to ensure no cluster sits idle while another queues work.
Real-Time observability and long-term FinOps auditing
You cannot govern or trust an AI data foundation that you cannot natively observe. This release expands Galaxy’s observability footprint with direct open integrations and long-term historical analysis.
Enterprise observability integrations
Direct OpenTelemetry Export to Datadog and AWS CloudWatch
Galaxy supports direct, push-based export of metrics, events, and traces into customer observability stacks via dedicated OpenTelemetry collectors. Platform teams gain real-time alerting on streaming ingestion freshness, query latency, and resource health inside their existing enterprise dashboards, backed by zero-trust, single-tenant OpenTelemetry collector isolation.
High-Scale query history and extended insights reporting
Moving query history storage off metadata databases onto a high-scale analytical store enables 30-plus day historical retention. Platform administrators can perform long-term FinOps auditing, cost tracking, and compliance analysis with sub-second report response times in the Galaxy UI.
A developer experience designed for speed and immediate value
The Query Editor is the landing page where data engineers and analysts spend the majority of their time. This release delivers a complete overhaul of the developer interface alongside guided onboarding workflows.
IDE and navigation enhancements
Redesigned Query Editor and Unified IDE
Features a modernized SQL editor with catalog-first navigation, cleaner tab management, flexible statement execution toggles, and enhanced autocompletion for deep catalog object hierarchies. Advanced tab overflow controls, tab pinning, and drag and drop reordering handle complex multi-query workflows at scale. Built on a unified codebase across Starburst Enterprise and Starburst Galaxy, feature updates and IDE enhancements land simultaneously across both environments.
Data-Centric navigation and enhanced filtering
Re-architects the Data Explorer sidebar away from cluster-first hierarchies toward a data-centric mental model focused directly on catalogs, schemas, tables, and data products. Column-level filtering and direct drag and drop column insertion streamline query composition.
Guided Galaxy onboarding
Streamlined onboarding flows eliminate early-stage friction by providing clear catalog-cluster linking, opinionated cluster sizing defaults, and guided setup paths, drastically reducing median Time-to-First-Query.
Enterprise interoperability and zero-downtime uptime
Production platforms demand continuous availability and seamless interoperability across existing enterprise governance investments.
Lifecycle management and open governance
Decoupled engine and Control Plane lifecycle
Trino engine version updates are now decoupled from Galaxy platform releases. Control plane updates no longer interrupt active Trino engine queries, enabling zero-downtime platform upgrades.
Expanded Unity Catalog and Delta Lake interoperability
Sustaining our commitment to open interoperability, Galaxy expands write support for Unity Catalog catalog-owned managed tables. Extended credential vending and REST-first Unity Catalog integration allow organizations to maintain Databricks Unity Catalog as a centralized governance layer while leveraging Starburst for multi-source, high-performance SQL analytics.
Engineering the AI foundation today
The breakthrough that unlocks enterprise AI in production was never going to be a larger model. It was always going to be a data foundation solid enough to trust. With Project AIDA, native AI Safety Guardrails, programmatic Data Quality APIs, editable metadata management, unified Iceberg management, global dynamic caching, sub-second log analytics, 2-3x higher query throughput, smart load balancing, native OpenTelemetry export, and a modern developer experience, Starburst Galaxy provides the managed, cloud-native foundation your data strategy demands.
Explore the Documentation and Get Started
- Starburst Galaxy Documentation
- Icehouse Console and Table Maintenance Guide
- Start your free Starburst Galaxy trial today



