Every agent is a loop over a model that calls tools. Union makes each turn of that loop a durable, containerized, replayable step — so a crash on turn seven resumes at turn seven, one tool can want an H100 while the next wants 200 MB, and six months later the run still explains itself.
Seven parts of the runtime under your agent, and where each one lives on this page.
Agent frameworks are good at planning the next step. They are less good at what happens when the worker running the agent dies on turn seven, when one tool needs a GPU and another needs 200 MB of RAM, or when someone asks what the agent actually did last Tuesday. That half is a runtime’s job.
The loop dies on turn seven.
Every completed turn and every tool result is checkpointed. The retry replays them from cache — no re-billed tokens, no re-run side effects — and resumes at the step that failed. The agent does not start over, and neither does your bill.
One tool needs an H100. The next needs 200 MB.
Each tool is an @env.task with its own image, resources, secrets, cache policy and retries. The model call stays in a small container; the fine-tune goes where the GPUs are. Fan-out is asyncio.gather(), and each branch gets its own hardware.
Nobody can say what it did last Tuesday.
Because the model decides the control flow at runtime, the only honest record is the one the run writes itself. @flyte.trace captures every model call, tool invocation and routing decision as a span with its inputs and outputs, and report=True renders the whole thing as a timeline you can reopen months later.
The model wrote rm -rf /.
Generated code is untrusted code — not because the model means harm, but because it has no way to know. Run it in a sandbox: Monty for orchestration logic, where imports, filesystem and network simply do not exist, or an ephemeral container when the code genuinely needs numpy and a disk.
How you build the agent and how you run it are orthogonal choices. Any loop below can later be put behind a schedule, a webhook or a chat UI without rewriting it — so start with whichever one gets you to a working agent fastest.
Union imposes no graph, no node type and no loop structure. Routing is if/else. Fan-out is asyncio.gather(). State and checkpointing are automatic. So the architectures are the ones you would have written anyway — they just happen to be durable.
Reason → Act → Observe, looping until the model stops calling tools. Use it for Single-agent tasks with tools: API calls, search, code execution.
@tool stacks on either. On a @flyte.trace function it stays in the agent’s own process and becomes a checkpointed span — cheap, and right for a lookup or a bit of parsing. On an @env.task it becomes a child action on the cluster with its own container, image, dependencies, resources, secrets, cache and retry policy — right for anything that needs a GPU, a big machine or a package the agent does not carry.
Either way the call is durable: its inputs and outputs are recorded, it gets its own row in the run timeline, and it replays from cache on a retry instead of running twice.
{name: tool} mappingrequires_approval=Truecall_handler — right-size it, then retry bigger on OOMRunning LLM-generated code without a sandbox means trusting the model never to make a mistake. Sandboxing removes the trust requirement. Which one you want depends on whether the model is writing control flow or computation.
Monty is a Rust-based sandboxed Python interpreter. It starts in microseconds and runs pure control flow — no imports, no filesystem, no network — dispatching the real work to container tasks through the Union controller. Dangerous operations are not restricted; they are structurally impossible.
Sandboxing ↗The parts of an agent that are nobody’s favourite to build, and that every agent eventually needs.
How the agent gets invoked is independent of how it was built. All three wrap the loop in a regular Union task, so every run is durable, retryable and observable — whichever door it came through.
The simplest deployment: the loop lives in an @env.task you invoke on demand from the CLI, a notebook or another service. report=True gives the run its timeline.
report=True renders the loop as a timeline: turns, tool calls, approvals, retries.AgentEvent — subscribe to stream progress anywhere.Every adapter exports tool and run_agent. tool stacks on @env.task, so a tool call becomes a containerized child action instead of a function call inside the agent process. run_agent drives the framework’s own loop from inside a durable parent task. Swap the import and the rest of the file stays as it is.
Put your loop in an @env.task and run it on Union. Nothing about the agent has to change to make it durable — that is the whole point — and it runs inside your own cloud, not ours.