Agents on Union

Kubernetes-native Agents. Durable by default.

Every agent is a loop over a model that calls tools. Union makes each turn of that loop a durable, containerized, replayable step — so a crash on turn seven resumes at turn seven, one tool can want an H100 while the next wants 200 MB, and six months later the run still explains itself.

Run any harness
Why this is a runtime problem

Frameworks decide what to do next. Something else has to record it.

Agent frameworks are good at planning the next step. They are less good at what happens when the worker running the agent dies on turn seven, when one tool needs a GPU and another needs 200 MB of RAM, or when someone asks what the agent actually did last Tuesday. That half is a runtime’s job.

The loop dies on turn seven.

Every completed turn and every tool result is checkpointed. The retry replays them from cache — no re-billed tokens, no re-run side effects — and resumes at the step that failed. The agent does not start over, and neither does your bill.

One tool needs an H100. The next needs 200 MB.

Each tool is an @env.task with its own image, resources, secrets, cache policy and retries. The model call stays in a small container; the fine-tune goes where the GPUs are. Fan-out is asyncio.gather(), and each branch gets its own hardware.

Nobody can say what it did last Tuesday.

Because the model decides the control flow at runtime, the only honest record is the one the run writes itself. @flyte.trace captures every model call, tool invocation and routing decision as a span with its inputs and outputs, and report=True renders the whole thing as a timeline you can reopen months later.

The model wrote rm -rf /.

Generated code is untrusted code — not because the model means harm, but because it has no way to know. Run it in a sandbox: Monty for orchestration logic, where imports, filesystem and network simply do not exist, or an ephemeral container when the code genuinely needs numpy and a disk.

Build

Four ways to build the loop. All of them are just tasks.

How you build the agent and how you run it are orthogonal choices. Any loop below can later be put behind a schedule, a webhook or a chat UI without rewriting it — so start with whichever one gets you to a working agent fastest.

Patterns

The patterns you already know, as ordinary Python

Union imposes no graph, no node type and no loop structure. Routing is if/else. Fan-out is asyncio.gather(). State and checkpointing are automatic. So the architectures are the ones you would have written anyway — they just happen to be durable.

Pattern:
react.py

Reason → Act → Observe, looping until the model stops calling tools. Use it for Single-agent tasks with tools: API calls, search, code execution.

Tools

A tool call can be a container or a durable function call

@tool stacks on either. On a @flyte.trace function it stays in the agent’s own process and becomes a checkpointed span — cheap, and right for a lookup or a bit of parsing. On an @env.task it becomes a child action on the cluster with its own container, image, dependencies, resources, secrets, cache and retry policy — right for anything that needs a GPU, a big machine or a package the agent does not carry.

Either way the call is durable: its inputs and outputs are recorded, it gets its own row in the run timeline, and it replays from cache on a retry instead of running twice.

→Rename or re-describe a tool for the model with a {name: tool} mapping
→Gate the irreversible ones behind requires_approval=True
→Intercept how a tool runs with call_handler — right-size it, then retry bigger on OOM
Tools reference ↗
Anything can be a tool
right_size.py call_handler
Sandboxing

Two sandboxes, because generated code has two shapes

Running LLM-generated code without a sandbox means trusting the model never to make a mistake. Sandboxing removes the trust requirement. Which one you want depends on whether the model is writing control flow or computation.

code_mode.py flyte.sandbox

Monty is a Rust-based sandboxed Python interpreter. It starts in microseconds and runs pure control flow — no imports, no filesystem, no network — dispatching the real work to container tasks through the Union controller. Dangerous operations are not restricted; they are structurally impossible.

Sandboxing ↗
Workflow sandboxCode sandbox
RuntimeMonty — a Rust-based Python interpreterAn ephemeral Docker container
StartupMicrosecondsSeconds — image build plus spin-up
CapabilitiesPure control flow. No imports, no I/O, no network.Any package, any library, full I/O
Security modelDangerous operations are impossible, not merely blockedIsolated container, discarded after one run
Reach for it whenThe model is writing orchestration that calls your tasksThe model is writing computation: ETL, tests, shell pipelines
The rest of the harness

Memory, MCP, approvals, and a built-in UI

The parts of an agent that are nobody’s favourite to build, and that every agent eventually needs.

Ship

One agent, three front doors

How the agent gets invoked is independent of how it was built. All three wrap the loop in a regular Union task, so every run is durable, retryable and observable — whichever door it came through.

agent.py flyte run

The simplest deployment: the loop lives in an @env.task you invoke on demand from the CLI, a notebook or another service. report=True gives the run its timeline.

Deploy an agent as a service ↗
Watching it run
→ Task reports. report=True renders the loop as a timeline: turns, tool calls, approvals, retries.
→ Grafana. Generations, tool calls, token usage and cost, grouped by run.
→ OpenTelemetry. Tasks and traced steps export as spans; a durable run arrives as one trace.
→ Events. Every step emits a typed AgentEvent — subscribe to stream progress anywhere.
Frameworks

Bring your favorite harness, run it durably

Every adapter exports tool and run_agent. tool stacks on @env.task, so a tool call becomes a containerized child action instead of a function call inside the agent process. run_agent drives the framework’s own loop from inside a durable parent task. Swap the import and the rest of the file stays as it is.

Tutorials

Twelve agents, end to end

All agent tutorials ↗

Start with the loop you already wrote.

Put your loop in an @env.task and run it on Union. Nothing about the agent has to change to make it durable — that is the whole point — and it runs inside your own cloud, not ours.

Get Started Read the agent docs ↗ pip install flyte