Inference on Union

Serve models and govern traffic inside your perimeter.

Union hosts OpenAI- and Anthropic-compatible endpoints on your own GPUs and puts every model your teams call behind one gateway. Easily manage external LLM providers, virtual keys, budgets, rate limits and guardrails.

Same Python API as the workflows that trained the model, same cloud account, same audit trail.

Why serving is a runtime problem

A model that works in a notebook is not a model in production.

The gap between a working checkpoint and a served one is almost entirely infrastructure: the idle GPU bill, the cold start, the key nobody can revoke, and the fact that every prompt your company writes is leaving the building. None of that is a modelling problem.

The GPU is idle most of the day.

An endpoint declares replicas=(0, 2) and a scaledown_after. Between bursts it holds zero replicas and costs nothing; the first request after that brings one back. The replica-count chart in the console is a sawtooth, and that sawtooth is the bill you did not pay.

A 27B checkpoint takes minutes to come up.

Most of that is copying weights to a disk you are about to read once. stream_model=True streams them from blob storage straight into GPU memory instead, and the image itself is cached on the node, so the cold path is the model load rather than the model download.

Nobody knows what the LLM spend is, or whose it is.

Provider keys get pasted into notebooks and CI and never rotated. The gateway inverts that: one endpoint, one key per team or agent, each with its own provider allowlist, monthly budget and rate limit. Spend is attributed to the key that caused it — including on models you host yourself, priced with numbers you supply.

Every prompt leaves your network.

Endpoints run on nodes in your own cloud account and are reachable only from inside it. The gateway runs there too, so the request, the completion and the audit record all stay on your side of the boundary — and when a call does go to a third-party provider, it is the one call you decided to allow.

Host the model

Deploy real-time model endpoints and dashboards.

The same Python API that runs your workflows serves long-running apps next to them. Native integrations wrap vLLM, SGLang, Ollama, FastAPI and Streamlit — declare the app, the image, the resources and the scaling, and Union handles the serving plumbing.

serve_llm.py flyteplugins-vllm

An OpenAI-compatible server in fourteen lines. stream_model=True pulls the weights straight from blob storage into GPU memory, so a cold replica does not wait on a full disk download first.

Native app integrations ↗
Govern the traffic

One OpenAI-compatible endpoint for every team and every provider.

The LLM Gateway is a Union app that sits in front of everything your teams call — the models you host, and the twenty upstream providers you buy from. Point any OpenAI-compatible SDK at it, authenticate with a virtual key, and address models as provider/model-id. GET /v1/models lists exactly what the calling key can reach.

Use it

No client change beyond a base URL and a key. Streaming, tool calls and the rest of the OpenAI surface pass straight through.

Throughput is a number, not a cluster exercise: one replica serves roughly 1,000 req/s, and min/max request rates set the floor that always runs and the burst ceiling above it.

Providers

Add an upstream with a Union secret holding its API key — the secret is mounted into the gateway pod, and the control plane never reads it. Restrict a provider to an allowlist of models, and give it a load-balancing weight against the others.

Virtual keys

Every team, agent and job gets its own key, budget and ceiling.

A virtual key is a scoped credential for calling the gateway. Each one carries the providers it may reach, a monthly budget, and a rate limit — so an agent that goes into a loop at 3 a.m. hits its own ceiling instead of your company’s invoice. Revoking a key is one row, not a key rotation across four repos.

YOUR OWN MODELS Custom Union models behind the same key

An endpoint you host on your own GPUs registers with the gateway like any other provider. Callers address it as union/qwen3-27b and never learn whether it is yours or a vendor’s — which is what makes swapping one for the other a config change.

BUDGETS You supply the cost

Upstream providers price themselves. For the models you host, you give the gateway the per-token cost — your amortised GPU rate, your internal chargeback number, whatever your finance team actually uses — and self-hosted traffic lands in the same spend column as everything else.

ATTRIBUTION Spend has a name on it

Requests, spend and keys-with-traffic are reported per key over the last 24 hours, and the logs carry the key on every request. “Who spent that” is a lookup rather than an investigation.

Guardrails

Two hooks around every call — and both of them are your code.

A guardrail you cannot read is a guardrail you cannot trust. Rather than a policy DSL, the gateway calls an ordinary HTTP app you deploy on Union: once before it forwards a request, once after the response comes back. Whatever you can write in Python, you can put in the path.

HookGetsCan
Pre-callThe request, the model, the virtual keyAllow it, deny it with a reason, rewrite it, or answer it from a cache of your own
Post-callThe completion and its token usage and costPass it through, replace it, redact it, or write the whole exchange to a sink you control
BothDeployed as Union appsScale, version and roll back like any other app — and run inside the same perimeter as the traffic
See it

The gateway is an app too, so it is watched like one.

Same log stream, same metric panes, same scaling events as any endpoint you deploy — because it is the same object. A routing change is a policy version you can point at, not a mystery.

Logs

Filter logsLive

Metrics

Point one client at one endpoint. See where the money goes.

Stand up the gateway, issue a virtual key, and move a single service onto it. Nothing in your application changes except a base URL — and for the first time the answer to “what are we spending on inference” is a number.

Get Started Read the serving docs ↗ pip install flyte