Multi-cluster scale-out study

Orchestration without Limits

Union scales out to run 200,000-action workflows at low latency, where single-cluster OSS Flyte cannot.

100 concurrent runs × 2,000-wide fan-out | identical task image · same client driver | 8 GiB OSS executor vs. Union's ScyllaDB fleet
AbstractFlyte v2 decomposes a run into many small per-action records rather than one big in-memory object — but v2 itself ships in two deployments: open-source Flyte, on a single cluster, and Union, which spreads its control and data planes across many. Both implement the same per-action model, so we ask whether decomposition alone delivers unbounded scale, or the surrounding architecture still sets the ceiling. We stress both with a swarm of K independent runs, each a 2,000-wide fan-out, ramped to 100 × 2,000 = 200,000 actions. Up to 20,000 actions both planes complete identically. At 200,000 actions they diverge sharply: Flyte OOM-kills its 8 GiB executor — its footprint tracks cumulative actions processed, not live count, and never releases — while Union completes all 100 runs in ~26 minutes. Beyond the reliability ceiling, Union also wins throughput outright: 1.5× on wide fan-out and 2.0× on sustained concurrency at 40,000 held tasks, with no runtime failure mode across any shape.

Read the full paper (PDF) ↗

200,000
cumulative actions Union completes; OSS Flyte OOMs here
26 min
wall-clock for Union to finish all 100 runs at 200k
~150,000
cumulative-action ceiling where OSS Flyte's executor OOMs
2.0×
faster on sustained concurrency at 40,000 held tasks
Key results

What the numbers show

Every chart below is drawn straight from the measured data in the paper — hover any bar or point for the exact value.

K runs × 2,000 actions — wall-clock (log scale)Fig. 1 · the swarm ramp
Flyte (OSS)Union
0 s 975 s 1,950 s 2,925 s 3,900 s 4k 102 s 59 s 10k 229 s 123 s 20k 440 s 185 s 50k 1,333 s 418 s 100k 3,617 s 772 s 200k OOM-KILLED 1,533 s

Every run completes on both planes through 100,000 actions. At 200,000 — the target scale — Flyte's 8 GiB executor OOM-kills and returns no runs, while Union completes all 100 in 1,533 s.

OSS Flyte executor memory vs. cumulative actions processedFig. 2 · cumulative-churn memory
Flyte (OSS) executor
0.0 GiB 2.3 GiB 4.5 GiB 6.8 GiB 9.0 GiB 20k 50k 100k 150k

Memory tracks cumulative actions processed, at ~54 MiB per 1,000 — not live count — because completed actions are never garbage-collected from the executor's informer cache. The line crosses the 8 GiB limit at ~150,000, exactly where the 200k swarm OOMs. Union stores action state in ScyllaDB, not a single process's heap, so no equivalent line exists for it.

Concurrent held tasks — wall-clockFig. 3 · held concurrency
Flyte (OSS)Union
0 s 205 s 410 s 615 s 820 s 1k 5k 10k 20k 40k 60k 80k OOM

Flyte's executor OOM-kills at 60,000 concurrently-held tasks — the live actions overflow the pod. Union holds 80,000 flat, already 2.0× ahead of Flyte at 40k and the only plane past it.

Flyte (OSS)  3,617 s

See it for yourself

Spin up a real Flyte cluster in minutes and run the same workloads on your own hardware. No infra to provision, no YAML to write first.