Train models, build datasets, and serve endpoints all inside your own cloud or datacenter, governed and owned by you. Your data never transits Union's infrastructure. Not proxied, not cached, not “encrypted in transit through us”. It just never leaves.
Self-hosted control plane is an enterprise plan, where nothing leaves your infrastructure. Explore Enterprise with us →
Four properties of a sovereign platform. Follow a card to the detail.
Most platforms ask you to trust their process. Union's isolation is topological, not behavioral: the data path physically does not pass through the control plane, and you can verify it — in the Helm charts, the tunnel config, and the Envoy filter chains.
The control plane stores run IDs, schedules, and artifact references. Inputs, outputs, logs, code bundles, and reports live in your object store and are served by a dataproxy inside your perimeter.
Your cluster initiates the tunnel. No inbound ports, no firewall changes, no VPN peering into your VPC. Every request is authenticated and authorized by an Envoy router that runs on your side of the line.
Data is encrypted with keys you hold in your KMS. Compromising Union's infrastructure would still not reach your data — an attacker would also need your cloud IAM and your keys.
Direct-to-DataPlane isn't a security tax. Reads and writes go straight to your object store, so the secure path has lower latency and lower egress cost than proxying through a vendor.
From the announcement: Zero-Trust Security Architecture
Every surface in the run view carries a lock — inputs, outputs, logs, reports, code. Hover over one and the console tells you exactly where the bytes came from: served by your data plane, through the Direct-to-DataPlane tunnel, never entering the control plane in any form.
The URL under the lock is your data plane's domain — the console fetches these bytes from your infrastructure, not Union's.
Sovereignty is worthless if it costs you capability. Union runs the full loop — data, training, evals, serving — on your GPUs, in your VPC, with the same event-wired automation as any managed platform.
Multi-node distributed training with ClusteredTaskEnvironment — torchrun across your own GPU nodes, checkpoints in your object store, OOM recovery mid-run.
Fan out synthesis and filtering across thousands of spot containers with flyte.map. Every element checkpointed; every dataset a versioned artifact with lineage.
Model endpoints, batch inference, and apps run on the same runtime that trained the model — behind your load balancer, reachable only from your network if you choose.
“References, never payloads” is a claim you can check asset by asset. Here is every thing a run touches, and which plane actually holds it.
The perimeter isn’t only the outer wall. Inside it, projects and domains are real boundaries: separate namespaces, separate service accounts, separate buckets and secrets, and network policy that keeps one team’s workloads off another team’s services.
Every action has a deterministic name, and the logs, metrics, errors, inputs, outputs, and cost of a step all hang off it. There is no correlation step, so the queue view, the cost rollup, and a task’s own report are all reading the same state the scheduler runs on.
Token-level clipped-ratio GRPO. The importance ratio comes from vLLM sampling-time logprobs against the adapter-disabled base, which is what makes the one-step-off-policy pipelined training sound.
Reports are published by the task itself, so “how did it do” has an answer in the same place as “did it finish”.
The teams already running these properties, and the number each of them reported.
Union is built on Flyte, the open-source AI runtime we create and maintain under the Linux Foundation AI & Data.
Owning where your AI runs is the first decision. These are the three that follow: a runtime that survives your infrastructure, a factory that makes the models, and a scheduler that shares the fleet.
Sovereignty is worthless if the run dies. Recover, fork, and replay any workload — and change the hardware inside an except block.
Explore the runtime →Own your data, your models, and your AI — the full training, evaluation, and serving loop running inside the perimeter above.
See the factory →Every cluster inside your perimeter, shared fairly: priority, quotas per team, and routing across clouds and accelerator classes.
See the scheduler →Bring one workload — training, data, or serving. We'll run it inside your cloud this week, and you can inspect exactly what the control plane sees: references, and nothing else.