Share

Linkedin iconFacebook iconTwitter icon

More deployment options

Imagine you are an engineer on call, looking at a failed CI run at 2am and wondering whether anyone has seen this failure before.

Now imagine you are a salesperson holding a customer’s security questionnaire, due Friday, with the answers scattered across a dozen documents.

Finally, imagine you are the person responsible for letting software answer both of them, with access to the systems that hold the truth. You have to decide what it may read, what it may change, who it is acting for, and how you find out when it is wrong.

That third person is who this post is about.

Two engineers started building Nebula, our internal agent platform, in July. Today 82 agents run on it, owned by teams across Starburst, and some of the people who built them do not write code for a living. In the rest of this post we will a) explain why we built a platform instead of a collection of agents, b) lay out the principles we designed to, c) walk through the stack and the primitives agents get for free, d) show what our pull request history says about where the work went, and e) tell you what broke and what we have not solved.

The thesis fits in a sentence. Without trust, an agent does more harm than good. Everything below is a consequence of taking that seriously.

Why a platform

By mid-year, teams here had built agents for on-call support, security operations and cost management. Each one worked. Each had also built its own runtime, deployment, identity and logging, because nothing existed to share. Nobody could answer the basic fleet questions: what is running, what can it touch, what does it cost, and who gets paged when one misbehaves.

Models had made agents cheap to write, so every team was about to build one, and reviewing bespoke agents one at a time does not scale to a company. A platform team’s job is leverage. Build the hard parts once, and make the safe route the easy one, a paved path teams choose because it is faster than building their own.

So we settled ownership first, and it fits on one line. We, central AI Infra, own the agent plane. All other teams at Starburst own the fleet. The plane is the shared layer every agent runs on, built once by one team. The fleet is the agents, built by everyone else. The platform owns everything that is not domain logic: Slack, model access, tool credentials, identity, deployment, scheduling and telemetry. An agent owns its prompt, its tools and its judgment.

The mental model is tenancy. Trino treats a connector as a managed entity that lives inside the engine, upgrades with it and is observed through it. We treat an agent the same way. Platform and agents share one repository and one deployment boundary, so a fix to the shared library reaches every agent in a single commit, with no SDK versions to pin and no upgrade backlog.

Design principles

Four principles guided the work. Each one is a way of buying trust.

Keep agents narrow. An agent is a small directory: a manifest, a factory, prompts and tests. The manifest declares identity, model, description, runtime, side effects, whether the agent accepts scheduled runs, and the grants it claims. It is kept short on purpose, so it reads as a statement of intent. Privileges such as a memory table, a shared-store prefix or knowledge-base access are set in the infrastructure config, next to the IAM that bounds them. Most agents are read-only or draft-only, and a human lands every change.

Make identity the platform’s job. A tool should never take “who is asking” as an argument, because a model can be talked into passing the wrong answer. The platform binds the asker’s identity to each turn, and tools read it from there.

Treat outside content as data, never as instruction. Stored objects are fenced by default, fetching is allowlisted and guarded against server-side request forgery, and a write guard cancels a Slack post unless the destination is one the user named or one a trusted tool returned.

Make the safe path the easy path. Credentials live at a gateway, tool access comes from named grants, and every call writes an audit record. An engineer gets all of it by describing an agent, and would have to work to get around it.

The stack

Dev interface. There is no CLI. The interface is Claude Code, working from runbooks checked into the repo. An engineer types /nebula new and answers a handful of questions about purpose, model tier, runtime, owner and tool grants. Claude writes the agent, runs it locally and fixes what fails, and the same entry point lists, runs, stops and modernizes agents. Local runs use the developer’s own cloud session, so no provider credential ever lands on a laptop. Platform knowledge lives in its own repository: scope, approach, findings, enforced vocabulary, a session log and work items. Every session loads it as context, which is how a new session picks up where the last one ended. The gap between an idea and a deployed, governed agent is hours.

Image depicting the dev interface for Starburst Nebula

Build and deploy. Every pull request runs tests, a linter, security scanning and an infrastructure preview. A merge to main deploys through GitHub Actions and Pulumi over OIDC, with no stored credentials. The pipeline discovers new stacks and rebuilds the agent registry from the manifests, so adding an agent is close to a single merge.

Runtime. Every agent ships as one container image with a two-route contract, POST /invocations and GET /ping. The same image runs locally, as a long-running service on ECS Fargate, or on Bedrock AgentCore as per-second microVM sessions that cost almost nothing when idle. The platform picks where each agent runs, based on what the agent’s manifest declares. An agent that needs to stay up for a long time, such as one that people talk to in Slack, runs on Fargate as a dedicated task that lives as long as the agent does. An agent that wakes on a schedule, does one job, and exits runs on AgentCore, which gives each run its own isolated machine and bills only for the seconds the agent is working. Agents that do not need a persistent task do not get one, which is where the cost saving comes from. Agents use Strands on Bedrock, with the model declared in the manifest.

The front door. People reach every agent through one Slack app. It resolves who is asking, answers discovery (“which agents can I talk to?”), and routes each thread up a ladder, cheapest first. An explicit to: prefix wins. Otherwise a thread already bound to an agent stays on it, which is what makes threads sticky. Otherwise one classifier call reads the agents’ registry descriptions, and if it is not confident, the request goes to a default agent instead of a guess. One shared front door means N agents need N+1 containers, not 2N, and one identity path and one audit trail instead of many.

Access: tools, identity and models

Access has three layers, and we keep them separate because they fail differently.

Tools come through named grants. A grant maps to a target on one managed tool gateway, an MCP gateway. The gateway holds the provider credential and the agent never does, so there is nothing in an agent’s code or environment to leak. We have 31 grants across nine systems, including GitHub, Atlassian, Datadog and Slack. A few conventions do most of the work:

  • Read and write are separate grants, and writes split further. Authoring a branch, opening a pull request and reviewing one are three grants, so an agent that only updates existing pull requests cannot create arbitrary branches. Merging is in none of them.
  • Allowlists are pinned in a catalog, and the catalog refuses to load if two grants on one target claim the same tool, so no edit can quietly widen an agent’s reach.
  • Some tools cannot be narrowed. A tool that dispatches by argument, running reads, writes or deletes depending on what you pass it, defeats an allowlist, because nobody can bound at review time what it will run at request time. We leave those out entirely.
  • A tenant boundary gets its own grant. We have one per Galaxy account, because an agent that wants product telemetry has no business holding the credential that reads revenue.
  • Claiming a grant is one reviewed line on a manifest in a pull request, like any other change.

Identity is rooted in Okta. The front door maps each Slack user to their Okta identity, each turn carries that identity in platform-bound context, and every tool call writes an audit record naming the user, the agent, the tool, the outcome and how the identity was established.

Models are declared in the manifest and reached through managed model access on Bedrock, so no agent holds a provider key. What we do not own yet is the path to the model. Every agent builds its own client, and the section on what broke explains what that cost us.

Primitives agents call

Beyond access, the platform exports a small set of things an agent calls from inside its tools. Each one came from a real agent needing it.

State. Invocations are stateless, which causes a problem the first time an agent needs to remember something. One of us hit it with a monitoring agent whose scheduled runs each start fresh. The first version kept its state in a secrets store and ran into a 64KB cap, no way to query, and an access policy that blocked the developer from their own secret. That friction became a primitive: one table per agent, opted in through the infrastructure config, exposed as plain get and set functions so the model never touches the database. We later added compare-and-swap, after two writers raced and handed out the same backlog id twice.

Shared store. Agents that build on one corpus need somewhere to put it. The shared object store gives each opting-in agent read or write access to named prefixes only. Because it crosses agent boundaries, reads are fenced by default.

Image depicting Starburst Nebula shared store.

Retrieval. An agent that names a knowledge base gets scoped retrieval on exactly that base. We grant retrieval and withhold generation, because letting a knowledge base invoke a model would reach outside the resource scope we pinned.

Files. Attachments move between Slack and tools through named extractors declared in the manifest, with a size ceiling and an isolated channel per invocation, so concurrent runs cannot see each other’s files. File content does not pass through the model.

Scheduling. Scheduling is a way of invoking an agent, not a separate kind of agent. A scheduler and an execution service call the same /invocations route on a timer. An agent opts in with one manifest field that defaults to off, and a scheduled turn carries its own identity class, so downstream records can tell a system-fired run from a user-initiated one.

Peer calls. An agent can call the peers its manifest names, through one tool per peer. For now those peers must be read-only agents on Fargate.

Streaming. Callers that ask get progress and heartbeat frames instead of waiting on one reply.

Observability. Since late September, every invocation emits metrics to Datadog from one place, the front-door listener, which is the only process that sees every turn on both runtimes. The scheduler emits the same metrics for scheduled turns. Agents write no instrumentation code. The platform attaches per-turn telemetry (tokens, cycles, stop reason, model and duration), and the listener prices it from a local table, because the model provider returns no cost. Invocations, outcomes, usage and estimated spend break out by agent, runtime, model and owner, with callers pseudonymized. Every log line carries the agent and, when a person asked, that person. A new agent appears on the dashboards the first time someone calls it.

What our pull request history says

Since the first pull request on July 15, 569 have merged. 55% are agent work and 24% are the platform. In the last 30 days, 230 merged and 40 of them touched the platform. A quarter of what we ship, three months in, is still the plane. We did not build the platform and finish it. We keep building it because the fleet keeps asking for more.

What broke

Approvals. We built a guard that lets a human approval through only when the approving speaker is the person who asked. The first version rejected every real approval, silently. The fix verifies the speaker against the turn’s own user and fails closed when the speaker cannot be verified. Failing closed is correct, and it also looks exactly like nothing happening, so you need a test that proves an approval can succeed.

Healthy and wrong. During a gateway outage, agents with dead tool grants kept answering health checks as healthy. We separated grant health from liveness. Three consecutive errors on a grant reads as degraded, and that signal stays out of the restart decision, so a provider outage does not become a restart loop.

Checks that passed on synthetic data. An agent working with sensitive records passed every check we wrote against synthetic inputs, then failed on the first live ones. A check that has only seen invented data has not been tested.

A shared service nobody could adopt. Before the platform existed, we stood up a model proxy, a service every model call could pass through so we could see cost and control routing in one place. No agent ever used it. Sending calls through it meant changing the few lines in each agent that make the model call, and every agent had its own copy of those lines. Nobody was going to edit dozens of agents to adopt a shared service. The lesson is that a shared service only gets adopted if the platform owns the path to it. We had standardized which model an agent declares in its manifest, but not how the call gets made. We are taking over that path now.

What we have not solved

Acting as the person. Today the audit log records who asked for every action. The system on the other end, GitHub or Jira for example, only knows the call came from Nebula, so it applies the platform’s permissions and records the platform’s name. We are building the path where the call is made as the person who asked, so the downstream system applies that person’s permissions and its own audit names a human.

Cost we can estimate but not prove. The model provider returns no cost per call, so we price each turn ourselves from token counts and a price table. We have not checked those estimates against the invoice. We do not allocate infrastructure cost per agent, there is no single health view across the portfolio, and nothing alerts when spend jumps, because the dashboards were set up by hand with no monitors.

Evals. An eval is a test that an agent still answers well on real inputs, run before it deploys. The building blocks exist inside one agent today: a harness that runs the real agent against the real model and real data, with every write sent to a sandbox and Slack blocked, and templates calibrated on real inputs before they ship. That is not a platform feature yet. It has to run on real inputs, come with every agent by default, and block a deploy instead of just reporting.

Drift. An agent is a prompt, a set of grants and a model, and each can change on its own. Someone edits the prompt, a grant is added, a model is upgraded. Over time an agent can stop doing what its one-line description says. The front door routes questions using those descriptions, so a drifted agent also receives the wrong questions. With more than eighty agents, we have no systematic way to catch that yet.

If agents are multiplying inside your company too, these are the five questions we would want answered for any of them. Who owns it. What can it touch. As to whom does it act. What does it cost. How do you turn it off. We can answer some cleanly and are still closing the rest. Cost is the one we cannot answer yet. We would like to compare notes with anyone further along.

Thank you to the various engineering teams and to every engineer who built an agent, including some who do not write code for a living.

Start for Free with Starburst Galaxy

Try our free trial today and see how you can improve your data performance.
Start Free