GPU-utilization study · FSG-26-04
Reuse or Reload
A GPU-utilization study of batch LLM inference: keep the worker warm, or spin a fresh container per call.
Qwen2.5-7B on one NVIDIA L4 | 500 GSM8K questions, ten 50-question chunks | measured with nvidia-smi every 2 s
Key results
What the numbers show
Every chart below is drawn straight from the measured data in the paper — hover any bar or point for the exact value.
See it for yourself
Spin up a real Flyte cluster in minutes and run the same workloads on your own hardware. No infra to provision, no YAML to write first.