AI engineering tip of the week: Run Claude Agent SDK agents on Flyte, with every tool as a task
Agent frameworks are good at deciding what to do next. They are less good at what happens when the process dies, when one tool needs a GPU and another needs 200 MB of RAM, or when someone asks what the agent actually did in a past run.
`flyteplugins-agents-claude` splits that work. You keep writing agents with the Claude Agent SDK. Flyte becomes the durable runtime underneath: every agent, sub-agent and tool call becomes a containerized durable Flyte task with its own resources, retries, and cache, and the whole conversation renders as a timeline in the task report.
Two decorators
The adapter exports two things, `tool` and `run_agent`. Stack `tool` on top of `@env.task` and the result is both a Flyte task and a tool Claude can call:
The docstring is the tool description Claude reads, so write it for the model. The input schema comes from the type hints through the Flyte type engine, which means `File`, `Dir`, dataclasses, and `Literal` enums all work as tool arguments.
Then `run_agent` drives the SDK's own loop from inside a task:
That task is the durable parent. `retries=3` makes the agent self-healing, and `report=True` gives you the timeline: assistant turns, each tool call with its arguments and result, turn count, wall-clock time, and a token breakdown with the SDK's cost estimate.
Why tools as tasks matters
In a normal agent process, a tool is a function call. Same machine, same memory limit, same failure domain. If the tool OOMs, the conversation dies with it.
Here, when Claude calls `lookup_order`, Flyte submits a child action. It gets its own container, its own resources, its own retries, and with `cache="auto"` an identical call returns from cache. The agent task itself stays small. A tool that needs a GPU can ask for one without the agent paying for it:
Tool calls are durable regardless of any other setting. The conversation itself is also protected: the adapter mirrors the Claude session onto a Flyte checkpoint, so a retried task resumes the conversation instead of restarting it. More on that in a future issue.
Fan out agents like any other task
An agent is a task, so you compose agents with ordinary Python:
Each request gets its own agent, its own container, and its own report. That is real distributed parallelism across workers, not asyncio inside one process.
Bring your own SDK options
Anything the Claude Agent SDK supports natively still works. Pass a `ClaudeAgentOptions` and the Flyte arguments layer on top:
Subagents, permissions, hooks, and session settings all pass through. If you set your own hooks, the Flyte observability hooks are merged in rather than replacing yours.
Same shape for many agent frameworks
The Claude adapter is one of many. OpenAI Agents SDK, Google ADK, LangGraph, LangChain, Deep Agents, CrewAI, Pydantic AI, Mistral, and Hermes each have a package built on the same core, and every one exports the same `tool` and `run_agent`:
Swap the import and the model name. The rest of the file stays as it is. A conformance test in CI keeps the surface identical across packages.
The `claude-agent-sdk` wheel bundles the native Claude Code runtime, so the image needs nothing beyond the pip install. Run locally with `--local` and an `ANTHROPIC_API_KEY` in your shell. Outside a task context the durability and report layers are no-ops, so the same file runs unchanged.
If you don't see a plugin for your agentic framework, Flyte can still run it. Wrap the agent, each sub-agent, and each tool-calling function with `@env.task` from a `TaskEnvironment`, and call them the way you would call any other Python function.
Every piece becomes a durable containerized task with its own resources, retries, caching, and a place in the run graph.
Full docs: https://www.union.ai/docs/v2/flyte/integrations/agents/claude-agent-sdk/
See what's happening in the Flyte Community:
Latest from the blog
- Inside Union: Resource-Aware Scheduling for Flyte 2.0 - Read on Union
- It's now super easy to start using Union.ai - Read on Union
- Union.ai Launches Self-Service Access on AWS Marketplace to Enable Sovereign AI - Read on Union
- Pandera 0.33: a CLI for validation without writing Python, and native PyArrow table support - Read on Union
- Inside Union: The Replay Log That Makes Flyte 2.0 Durable - Read on Union
- We Ran Multi-Node GRPO on 8 GPUs and the Trainer Cost Us Nothing - Read on Union
Recent talks & recordings
- Union: AI Infrastructure Made Easier - Watch on YouTube
- Flyte 2: The Durable Runtime Built for AI - Watch on YouTube
- When the Pipeline Breaks: Building ML Infrastructure for Biotech R&D | Session 1 - Watch on YouTube
- LLM fine-tuning with GRPO - Watch on YouTube
Upcoming events
- [Seattle] Build Your First Model Factory - Hacknight Seattle | Oct 22nd - RSVP on Luma
- [SF] Own Your AI: Build Your First Model Factory - Hack Night SF | Nov 3 - RSVP on Luma
Releases & updates
- Flyte 2 Is Generally Available: The Durable, Open-Source AI Runtime - Read on Union.ai
<div class="button-group is-center"><a class="button" target="_blank" rel="noopener noreferrer" href="https://www.union.ai/docs/v2/flyte/user-guide/run-modes/running-devbox/">Download Devbox</a></div>
From the community
- AI Book Club: Grokking Deep Reinforcement Learning - RSVP on Luma
That's all for this week! - Sage Elliott




