Run Flyte tasks as Slurm jobs on an existing HPC or GPU cluster, submitted over SSH.
Slurm
The Slurm plugin lets you run Flyte tasks as jobs on an existing Slurm cluster — an on-premise HPC installation, a cloud GPU cluster, or one managed by an operator such as Soperator. Jobs are submitted over SSH to a login node, so no Flyte component runs on the cluster and Slurm itself needs no reconfiguration. The plugin assumes nothing about which cloud the cluster runs in or where the run’s object storage lives.
What the cluster does need depends on the task type. A native slurm task runs your container image, so the compute nodes need Pyxis/Enroot or Apptainer and credentials for the run’s object storage, and requesting GPUs needs GRES configured. A slurm_script task needs none of that. The connector handles submission, state polling, cancellation and log retrieval.
The plugin supports:
- Running Python tasks on Slurm with the same typed inputs and outputs they would have on Kubernetes
- Running existing
sbatchscripts, including multi-node jobs - Full
sbatchscheduling options, either as first-class fields or passed through - Containerized execution through Pyxis/Enroot or Apptainer
- Slurm state mapping, so a queued job is not counted as running and a preempted job is retried
Installation
pip install flyteplugins-slurmThe connector must also be installed in the flyteconnector image of your data plane. See Deployment.
Task types
The plugin provides two task types, served by one connector.
| Task type | What is submitted | Typed I/O | Caching | Multi-node |
|---|---|---|---|---|
slurm |
The task’s own container image and the Flyte entrypoint, via Pyxis or Apptainer | Yes | Yes | No |
slurm_script |
A user-supplied sbatch script, with #SBATCH directives and input/output exports prepended |
File and Dir, declared |
Yes | Yes |
Prefer slurm for anything that can be containerized and runs as a single process. Use slurm_script when a script cannot be converted, or when you need gang-scheduled multi-node execution.
Quick start
Create a Slurm configuration and pass it as plugin_config to a TaskEnvironment:
import flyte
from flyte.io import File
from flyteplugins.slurm import Slurm
slurm_env = flyte.TaskEnvironment(
name="train",
plugin_config=Slurm(
partition="main",
nodes=1,
gres="gpu:8",
time_limit="4:00:00",
# Credentials for the run's object storage, mounted rather than passed in `env`.
container_mounts=["/home/flyte/.cloud:/etc/cloud:ro"],
env={"AWS_SHARED_CREDENTIALS_FILE": "/etc/cloud/credentials"},
),
image=flyte.Image.from_debian_base().with_pip_packages("flyteplugins-slurm"),
)
@slurm_env.task(cache="auto", retries=2)
async def train(steps: int = 1000) -> File:
...Remove plugin_config and the same task runs as a Kubernetes pod with no other changes. That is also the quickest way to tell a Slurm problem apart from a task problem.
resources on a Slurm task environmentThe allocation comes from the Slurm configuration and is granted by Slurm, not by Kubernetes. Setting resources raises an error. Use cpus_per_task, mem, gres or gpus_per_node instead.
Running an existing sbatch script
SlurmScriptTask submits an existing script. Scalar inputs are exported as FLYTE_INPUT_<NAME>:
import flyte
from flyteplugins.slurm import Slurm, SlurmScriptTask
train = SlurmScriptTask(
name="train",
script=open("train.sbatch").read(),
plugin_config=Slurm(partition="main", nodes=4, time_limit="8:00:00"),
inputs={"epochs": int},
)
env = flyte.TaskEnvironment.from_task("legacy-train", train)The plugin makes three changes to the script before submitting it:
- The script’s own leading
#SBATCHdirectives are moved to the top, followed by the plugin’s directives.sbatchtakes the last value for a duplicate option, so the plugin’s settings win on conflict and all other options are kept. exportlines for inputs and outputs are added after the directives.- A leading shebang is dropped.
Keep your #SBATCH directives above the first executable line, because sbatch stops reading directives there. Because the script drives srun itself, this is the task type to use for multi-node work.
A task has to be attached to a TaskEnvironment before it can be serialized. flyte.TaskEnvironment.from_task does that for a standalone task.
str, int, float and bool arrive as FLYTE_INPUT_<NAME>, and File/Dir as their URI for the script to fetch with its own tooling. Any other type fails the task at submission, because an environment variable cannot carry it. Pass a URI as a str instead.
Outputs from a script task
A script cannot write Flyte’s output format, so a script task returns only the outputs you declare. Without declared outputs, the task returns nothing.
Declare what the script will write
train = SlurmScriptTask(
name="train",
script=open("train.sbatch").read(),
plugin_config=Slurm(partition="main"),
inputs={"epochs": int},
outputs={"model": File, "shards": Dir},
)Only File and Dir may be declared. Any other type is rejected when the task is defined. A native slurm task supports all types; see Output types.
This happens even if the script exits with 0, so a downstream task never receives a URI that points to nothing.
Write to the destination the script is given
Each declared output arrives as FLYTE_OUTPUT_<NAME>, upper-cased. By default that is a local path, so writing the output is a cp:
python train.py --epochs "$FLYTE_INPUT_EPOCHS" --out ./model.pt
cp ./model.pt "$FLYTE_OUTPUT_MODEL"Every $FLYTE_OUTPUT_* the script uses must be declared in outputs, and every declared output must be written. Both are checked:
| The script has | outputs has |
What happens |
|---|---|---|
$FLYTE_OUTPUT_MODEL |
{"model": File} |
The output is recorded and a downstream task can read it |
$FLYTE_OUTPUT_MODEL |
nothing, or another name | ValueError when the task is defined |
| nothing | {"model": File} |
The task fails when the job finishes, even on exit 0 |
The definition-time check only sees $NAME and ${NAME} expansions in the script. A mention in a comment does not count, and a variable name built at run time is not checked.
When the job succeeds, the connector checks that every destination exists and records it as the declared File or Dir.
Choose who uploads
output_upload decides who moves the output bytes to object storage:
"connector" (default) |
"job" |
|
|---|---|---|
FLYTE_OUTPUT_<NAME> holds |
a local path | the object-storage URI |
| Who uploads | the connector, after the job | the script, during the job |
| The compute node needs | nothing | a client and credentials for the store |
| The bytes travel | node → connector → storage | node → storage |
| Size limit | 100 MiB by default | none |
"connector" suits small outputs such as metrics, summaries, configs and small models. The node needs no upload tool and no credentials.
The upload runs in the background after the Slurm job finishes. The task stays in RUNNING, with a message naming the outputs being moved, for a poll or two after the job ends. This is expected.
The task fails when an output is over the limit. The limit is only checked after the job has run, so the job’s work is lost. Use output_upload="job" for any output that might be large.
Operators can change the limit with FLYTE_SLURM_CONNECTOR_UPLOAD_MAX_BYTES on the connector deployment. It takes a byte count or a size with a suffix (500MB, 2GB, 512MiB), or 0 for no limit.
"job" has no size limit, because the bytes go straight from the compute node to storage:
train = SlurmScriptTask(..., outputs={"model": File}, output_upload="job")# S3, and S3-compatible stores (MinIO, R2, Ceph) with --endpoint-url
aws s3 cp ./model.pt "$FLYTE_OUTPUT_MODEL"
# Google Cloud Storage
gcloud storage cp ./model.pt "$FLYTE_OUTPUT_MODEL"
# Azure Blob
azcopy copy ./model.pt "$FLYTE_OUTPUT_MODEL"
# A directory output: copy the tree
aws s3 cp --recursive ./checkpoints "$FLYTE_OUTPUT_CHECKPOINTS"A script task runs directly on the node, not in a container, so the available tools depend on the site. Check on a compute node, not the login node, since they may differ:
srun --ntasks=1 bash -c 'command -v aws rclone gcloud azcopy'rclone is common on HPC clusters and needs no configured remote if you give it the backend inline. It takes a bucket path rather than a URL, so strip the scheme:
GCS=":gcs,service_account_file=$HOME/.gcp/sa.json,bucket_policy_only=true:"
rclone copyto ./model.pt "${GCS}${FLYTE_OUTPUT_MODEL#gs://}"
S3=":s3,provider=AWS,env_auth=true:"
rclone copyto ./model.pt "${S3}${FLYTE_OUTPUT_MODEL#s3://}"bucket_policy_only=true is required on a uniform-access GCS bucketWithout it, rclone sets a per-object ACL, and a bucket with uniform bucket-level access (the default for new buckets) rejects the write with Error 400: Cannot insert legacy ACL for an object. The job fails after doing its work.
Mount credentials the same way as for a native task. See Object storage from inside the job.
slurm tasksA native task runs the Flyte entrypoint inside the job, which writes its outputs to object storage directly. There is no upload mode to choose and no size limit.
Configuration
Scheduling
These fields map one-to-one onto sbatch options.
| Parameter | Type | Description |
|---|---|---|
partition |
str |
Partition to submit to |
nodes |
int |
Number of nodes to allocate |
ntasks |
int |
Number of tasks (--ntasks) |
cpus_per_task |
int |
CPUs per task |
gres |
str |
Generic resources, for example "gpu:8". Requires GRES configured on the cluster; see the warning below |
gpus_per_node |
int or str |
GPUs per node, for example 8 or "h100:8" |
mem |
str |
Memory per node, for example "64G" |
time_limit |
str |
Wall-clock limit in Slurm format, for example "4:00:00" |
account |
str |
Account to charge |
qos |
str |
Quality of service |
reservation |
str |
Reservation name |
constraint |
str |
Node feature constraint |
sbatch_options |
Dict[str, Any] |
Any other sbatch option, as --<key>=<value>. True renders a bare flag. Overrides the first-class fields on conflict. The plugin sets job-name, output, error and chdir itself, so these and wrap, uid and gid are rejected |
Container and execution
| Parameter | Type | Description |
|---|---|---|
container_runtime |
str |
How the image is launched: "pyxis" (default) or "apptainer". See Container runtimes |
container_image |
str |
Override the image given to the runtime, for example a pre-imported .sqsh or .sif path. Defaults to the task’s image |
container_mounts |
List[str] |
Bind mounts as src:dst[:ro], for example ["/data:/data"] |
container_workdir |
str |
Working directory inside the container |
container_args |
List[str] |
Extra arguments for the container runtime, for example ["--rocm"] for AMD GPUs under Apptainer |
modules |
List[str] |
Environment modules to module load before the job runs, for example ["apptainer"] |
srun_args |
List[str] |
Extra arguments inserted before the command on the srun line |
env |
Dict[str, str] |
Environment variables exported into the job |
working_dir |
str |
Directory for generated scripts and logs. Relative paths are under the SSH user’s home. Defaults to .flyte/jobs |
gres and gpus_per_node require GRES on the clusterGPU scheduling must be enabled on each cluster: the controller needs GresTypes set and each node needs its own Gres entry. On a cluster without them, any GPU request is rejected at submission and the task never starts:
sbatch: error: Invalid generic resource (gres) specification.Check what a cluster offers before requesting it:
sinfo -N -o "%N %G" # per-node GRES; "(null)" means none configuredenvValues in env are written into the generated sbatch script in plain text, and that script stays on the cluster filesystem. Mount credentials from the cluster’s shared filesystem and reference the path instead.
Connection
Each of these can be set per task, or once for the whole cluster on the connector deployment.
| Parameter | Connector environment variable | Description |
|---|---|---|
host |
FLYTE_SLURM_HOST |
Login node hostname |
port |
FLYTE_SLURM_PORT |
SSH port, default 22 |
username |
FLYTE_SLURM_USERNAME |
SSH user that jobs are submitted as |
ssh_private_key |
FLYTE_SLURM_SSH_PRIVATE_KEY |
Name of the Flyte secret holding the SSH private key |
known_hosts |
FLYTE_SLURM_KNOWN_HOSTS |
Path to a known_hosts file on the connector, for host-key verification |
known_hosts_secret |
— | Name of a Flyte secret holding the known_hosts entries. Needs no file mounted on the connector |
| — | FLYTE_SLURM_SKIP_HOST_KEY_VERIFICATION |
Disable host-key verification. Development only |
| — | FLYTE_SLURM_WORKING_DIR |
Cluster-wide default for working_dir |
| — | FLYTE_SLURM_CONNECTOR_UPLOAD_MAX_BYTES |
Largest script-task output the connector will upload; default 100MiB, 0 for no limit |
Setting the connection once on the connector is usually best: tasks then carry only scheduling options and stay portable.
For host, port, username and known_hosts, a value set on the connector wins over the task’s value. This stops a task from sending the deployment’s SSH key to a different host. A task’s value is used only when the connector sets none. FLYTE_SLURM_WORKING_DIR is the exception: a task’s working_dir overrides it.
Container runtimes
A native slurm task runs your image on the compute node, which needs a container runtime on the cluster. container_runtime selects it:
pyxis (default) |
apptainer |
|
|---|---|---|
| How it launches | flags on srun |
apptainer exec in the job |
| Image reference | ghcr.io#org/img:tag |
docker://ghcr.io/org/img:tag |
container_mounts becomes |
--container-mounts |
--bind |
container_workdir becomes |
--container-workdir |
--pwd |
| Local image | .sqsh path |
.sif path |
| GPUs | Enroot exposes them automatically | --nv is added when the job requests GPUs |
Slurm(partition="main", container_runtime="apptainer")Pyxis is common on GPU clusters built on NVIDIA’s stack; Apptainer is more common at traditional HPC sites. The rest of the job is the same for both, so a task moves between clusters by changing this one field. Check which runtime the cluster has:
scontrol show config | grep -i plugstack # Pyxis: look for spank_pyxis.so
command -v apptainer # ApptainerIf apptainer is only available through environment modules, add it to modules. A cluster with neither runtime cannot run native tasks; use slurm_script instead. An unknown container_runtime value is rejected when the task is defined.
Container images
The plugin converts the task’s image reference into the form the runtime expects:
| Input | Pyxis | Apptainer |
|---|---|---|
ghcr.io/myorg/train:v1 |
ghcr.io#myorg/train:v1 |
docker://ghcr.io/myorg/train:v1 |
python:3.12-slim |
unchanged | docker://python:3.12-slim |
/jail/images/train.sqsh or .sif |
unchanged | unchanged |
On clusters that pre-import images to a shared filesystem, point at the image file directly and skip the registry pull:
Slurm(container_image="/jail/images/train.sqsh", partition="main")Images are pulled on the compute nodes, as the submitting user. Your local Docker login and Kubernetes imagePullSecrets do not apply. For a private registry:
-
Pyxis: add an entry to
~/.config/enroot/.credentials:machine ghcr.io login <username> password <token> -
Apptainer: run
apptainer registry login --username <username> docker://ghcr.io, or setAPPTAINER_DOCKER_USERNAMEandAPPTAINER_DOCKER_PASSWORD.
Without credentials, the pull is anonymous and fails with 401 Unauthorized.
Object storage from inside the job
A slurm task runs the Flyte entrypoint inside the job, which reads inputs and writes outputs to the run’s object storage. Compute nodes therefore need network access to that storage and credentials for it.
The credentials depend on the store:
| Store | Variable the job needs |
|---|---|
| S3, and S3-compatible (MinIO, R2, Nebius, …) | AWS_SHARED_CREDENTIALS_FILE, or the usual AWS_* pair |
| Google Cloud Storage | GOOGLE_APPLICATION_CREDENTIALS |
| Azure Blob | AZURE_STORAGE_* |
Mount the credential file from the cluster’s shared filesystem with container_mounts and set its path in env, as in the quick start. Do not put the secret itself in env.
Outputs land in the same place they would for a Kubernetes pod task, so a Slurm task can pass results to a task running elsewhere:
slurm_env = flyte.TaskEnvironment(
name="train",
plugin_config=Slurm(partition="main", gres="gpu:8"),
image=image,
)
k8s_env = flyte.TaskEnvironment(name="pipeline", image=image, depends_on=[slurm_env])
@k8s_env.task
async def prepare(rows: int) -> Dir: ...
@slurm_env.task
async def train(dataset: Dir) -> File: ...
@k8s_env.task
async def evaluate(model: File) -> dict[str, str]: ...
@k8s_env.task
async def pipeline() -> dict[str, str]:
return await evaluate(await train(await prepare(1000)))File or Dir, not a cluster filesystem pathReturning a path such as "/data/model.pt" as a str passes type checking, then fails when a task on another cluster opens it, because the Slurm cluster’s filesystem does not exist there. Return flyte.io.File or flyte.io.Dir so the contents are uploaded. Plain paths work only between tasks that share a filesystem.
For data read repeatedly, such as a training set read every epoch, copy it to the Slurm cluster’s shared filesystem once and pass a path on the cluster. Reading it from object storage on every epoch is slow and expensive.
Outputs and caching
Output types
slurm |
slurm_script |
|
|---|---|---|
Scalars — int, float, str, bool, datetime, date, timedelta, None |
Yes | No |
Containers — list, dict, Union, Optional, Enum |
Yes | No |
| Structured — dataclass, Pydantic model, protobuf | Yes | No |
Offloaded — File, Dir, DataFrame |
Yes | File and Dir, declared |
| Anything else, via pickle | Yes | No |
Several outputs as a tuple |
Yes | Yes, all declared |
A native slurm task runs the same entrypoint as a Kubernetes pod task, so its outputs work exactly as they do on Kubernetes. A slurm_script task returns only declared File and Dir outputs, as described in Outputs from a script task.
Because the job does its own I/O, return large values as a File or Dir rather than directly. Inline inputs and outputs are capped by max_inline_io_bytes.
Caching
Both task types cache with cache="auto":
train = SlurmScriptTask(name="train", script=SCRIPT, outputs={"model": File}, cache="auto")A cache hit restores the declared outputs and does not submit the job.
For a native task, the cache version comes from the task function’s source, as for any Python task. A script task has no function, so the plugin computes the version from the script and its configuration instead:
| Change | Cache |
|---|---|
| The script body | Invalidated |
| A declared output added, removed, or retyped | Invalidated |
Scheduling and container config — partition, nodes, time_limit, gres, sbatch_options, … |
Invalidated |
host, port, username, ssh_private_key, known_hosts |
Reused |
output_upload |
Reused |
Moving to a new login node, rotating the SSH secret or changing who uploads does not change what the job computes, so the cache is kept.
An explicit Cache(behavior="override", version_override=...) is used as-is; the plugin only computes a version when the behavior is "auto". cache="disable" turns caching off.
A cache hit skips the job and restores the outputs. With no outputs declared, a hit just skips the job, which is rarely what you want from a script that runs for its side effects. Declare outputs, or set cache="disable".
Job state mapping
Only the first word of the Slurm state is matched, so CANCELLED by 1234 is treated as CANCELLED. An unrecognized state is logged and treated as running.
| Slurm state | Flyte phase | Notes |
|---|---|---|
PENDING, CONFIGURING, REQUEUED, SUSPENDED |
QUEUED |
Time waiting for an allocation does not count as running |
RUNNING, COMPLETING |
RUNNING |
— |
COMPLETED |
SUCCEEDED |
Script tasks with output_upload="connector" stay RUNNING until the connector has uploaded their outputs |
FAILED, NODE_FAIL, OUT_OF_MEMORY, TIMEOUT, DEADLINE, BOOT_FAIL, SPECIAL_EXIT, REVOKED |
FAILED |
The message includes Slurm’s reason and the end of stderr |
PREEMPTED |
RETRYABLE_FAILED |
Uses a retry instead of failing the run, but only if the task sets retries (default 0) |
CANCELLED |
ABORTED |
Aborting the Flyte run runs scancel |
State is polled with squeue, falling back to sacct for jobs that have left the queue. Make sure accounting works for the submitting user. Without it, the connector falls back to scontrol; see Known gaps.
Inspecting a job
Every job leaves three files on the login node under working_dir, named flyte-<task>-<8 hex>:
ls -t ~/.flyte/jobs | head
cat ~/.flyte/jobs/<job>.sbatch # exactly what was submitted
tail ~/.flyte/jobs/<job>.out # stdout: a successful job's output
tail ~/.flyte/jobs/<job>.err # stderr: usually where a failure explains itselfThe generated .sbatch file is a plain script. Reading it answers most questions, and running it by hand with sbatch tells a plugin problem apart from a cluster problem. Both log paths are also shown in the task’s status message in the UI. Nothing deletes these files, so prune them on a schedule that suits the site.
A Slurm task runs on a Slurm node, not in a pod. An empty pod list for the action is expected; check sacct on the cluster instead.
Deployment
Connector image
Add flyteplugins-slurm to the flyteconnector image:
FROM ghcr.io/flyteorg/flyte-connectors:<tag matching your data plane>
RUN pip install flyteplugins-slurmPush it to a registry the data plane can pull from.
Data plane values
Create the secret holding the SSH key and the known_hosts file:
kubectl -n <namespace> create secret generic slurm-login \
--from-file=ssh-privatekey=./slurm-connector \
--from-file=known_hosts=./known_hostsThen point the connector at your image and give it the connection:
flyteconnector:
# Also enables the connector endpoint in the leaseworker's config.
enabled: true
image:
repository: <your-registry>/slurm-connector
tag: <your-tag>
additionalEnvs:
- { name: FLYTE_SLURM_HOST, value: "<login-host>" }
- { name: FLYTE_SLURM_USERNAME, value: "flyte" }
- name: FLYTE_SLURM_SSH_PRIVATE_KEY
valueFrom:
secretKeyRef: { name: slurm-login, key: ssh-privatekey }
- { name: FLYTE_SLURM_KNOWN_HOSTS, value: /etc/slurm-login/known_hosts }
# These two take a map, not a list. See the warning below.
additionalVolumeMounts:
volumeMounts:
- { name: slurm-login, mountPath: /etc/slurm-login, readOnly: true }
additionalVolumes:
volumes:
- name: slurm-login
secret:
secretName: slurm-login
items: [{ key: known_hosts, path: known_hosts }]FLYTE_SLURM_SSH_PRIVATE_KEY holds the key’s contents, so it comes from a secretKeyRef. FLYTE_SLURM_KNOWN_HOSTS is a path, so its file is mounted. To avoid the mount, use known_hosts_secret on the task instead.
additionalVolumes and additionalVolumeMounts take a map, not a listThe chart inserts these values into the pod spec without its own volumes: / volumeMounts: key, so your value must include it. A plain list, which the chart’s [] default suggests, renders invalid YAML and fails the upgrade:
YAML parse error on dataplane/templates/flyteconnector/deployment.yaml:
error converting YAML to JSON: yaml: did not find expected keyadditionalEnvs takes a plain list. Render the template before upgrading:
helm template t <chart> -s templates/flyteconnector/deployment.yaml -f values-slurm.yamlOn upgrade, Helm also prints warning: destination for dataplane.flyteconnector.additionalVolumeMounts is a table. Ignoring non-table value ([]). This is expected.
The connector’s own object-storage access
With output_upload="connector" (the default for script tasks), the connector pod writes to the run’s output prefix. By default the flyteconnector service account does not have the data plane’s cloud identity, unlike union-system, webhook and dataproxy. Without it, every upload fails after the job has succeeded:
The operation lacked the necessary privileges to complete for path
metadata/v2/<org>/<project>/<domain>/<run>/<action>/0/<output>:
403 Forbidden ... Caller does not have storage.objects.create accessAnnotate the service account with the identity that can write the metadata bucket, the same one as in global.BACKEND_IAM_ROLE_ARN:
flyteconnector:
serviceAccount:
annotations:
# GCP
iam.gke.io/gcp-service-account: union-system@<project>.iam.gserviceaccount.com
# AWS
# eks.amazonaws.com/role-arn: arn:aws:iam::<account>:role/<backend-role>On GKE, also add the workload identity binding, or the annotation has no effect:
gcloud iam service-accounts add-iam-policy-binding \
union-system@<project>.iam.gserviceaccount.com \
--role roles/iam.workloadIdentityUser \
--member "serviceAccount:<project>.svc.id.goog[<namespace>/flyteconnector]"Then restart the deployment so the pods pick up the new identity. This is not needed for native slurm tasks or for script tasks with output_upload="job", because the job uploads with its own credentials.
Verifying the connector
Task types are discovered at runtime from the connector’s metadata service, so there is no task-type routing to configure. Confirm the connector advertises them:
kubectl -n <namespace> logs deploy/flyteconnector | grep -A6 "Connector Metadata"Credentials
The SSH private key is a Flyte secret named by ssh_private_key, or is set cluster-wide as FLYTE_SLURM_SSH_PRIVATE_KEY on the flyteconnector deployment. Provide known_hosts entries for host-key verification, either as a file (known_hosts) or a secret (known_hosts_secret). FLYTE_SLURM_SKIP_HOST_KEY_VERIFICATION on the connector disables verification for development; it cannot be turned off from a task.
Network
The connector must reach the login node on its SSH port. Restrict the login node’s allowed source ranges to the connector’s egress addresses, and test from the connector pod, not from your workstation:
kubectl -n <namespace> exec deploy/flyteconnector -- \
python -c "import socket; s=socket.socket(); s.settimeout(10); \
s.connect(('<login-host>', 22)); print(s.recv(64))"The connector keeps one SSH connection per cluster and reuses it across jobs, reconnecting if it drops.
Known gaps
Execution
- No multi-node execution for
slurmtasks. The native task runssrun --nodes=1 --ntasks=1, so the Flyte entrypoint runs once even when the allocation spans several nodes. Use aslurm_scripttask, which runssrunormpirunitself, for distributed work. - Only Pyxis and Apptainer are supported. A cluster with neither cannot run native tasks; use
slurm_scriptinstead.
Data and I/O
- Script task outputs are
FileandDironly. See Outputs from a script task. - Script task inputs are scalars and URIs only. See Running an existing sbatch script.
- The connector uploads at most 100 MiB per output by default. Raise
FLYTE_SLURM_CONNECTOR_UPLOAD_MAX_BYTES, or useoutput_upload="job". - No clickable log links. Job stdout and stderr are files on the login node, so their paths are shown in the task’s status message. Live logs are streamed through the connector.
Operations
- SSH transport only.
slurmrestdis not supported. - One identity. Every job runs as the configured SSH user, so the cluster attributes all work to that account, whoever launched the run.
- Status is polled per job. The connector runs one
squeueper job on each poll. - Job files are not cleaned up. See Inspecting a job.
- Clusters without accounting. Without
sacct, a finished job is looked up withscontrol, which only keeps it forMinJobAgeseconds. A job that finishes and ages out between two polls cannot be resolved.
API reference
See the Slurm API reference for full details.