Skip to content

Run on a Cluster#

This guide explains how to execute oxo-flow workflows on HPC clusters using SLURM, PBS, SGE, and LSF backends.


Overview#

oxo-flow's cluster module translates each rule into a cluster job submission. Resource requirements declared in the .oxoflow file (threads, memory, gpu, time_limit) are mapped to the appropriate scheduler directives. The disk field is not mapped to any scheduler directive — it only produces a local warning during oxo-flow run.

There are two ways to reach a scheduler, and they suit different jobs:

What it does Use when
run --profile <NAME> Submits, tracks jobs to completion, and updates the checkpoint Normal execution — you want the workflow to run
cluster submit Writes job scripts and stops You want to inspect or hand-edit scripts before submitting

run --profile is the everyday path; it inherits everything run does — wildcard expansion, checkpoint/resume, --samples, --rerun, config-change invalidation. See Cluster submission for the [cluster] profile block. cluster submit remains the escape hatch and is documented below.

Environment wrapping is applied automatically — conda, mamba, docker, singularity/apptainer, pixi, venv, and modules environments are wrapped in the generated scripts. Each rule applies one backend (the first declared in the resolver order mamba → conda → pixi → docker → singularity → venv → modules), so declaring e.g. both singularity and modules fails loudly instead of silently dropping the module loads.


Supported Schedulers#

Scheduler Status Directive prefix
SLURM Supported #SBATCH
PBS/Torque Supported #PBS
SGE Supported #$
LSF Supported #BSUB

Declaring Resources#

Set resource requirements per rule:

[[rules]]
name = "align"
input = ["{sample}_R1.fastq.gz"]
output = ["aligned/{sample}.bam"]
environment = { singularity = "docker://biocontainers/bwa:0.7.17" }
shell = "bwa mem -t {threads} ref.fa {input} | samtools sort -o {output}"

[rules.resources]
threads = 16
memory = "32G"
time_limit = "24h"

Note on gpu: gpu = 0 means "no GPUs" — the engine omits the GPU directive entirely, exactly as if the field were absent. Only declare it when you need at least one GPU (gpu = 1 or higher).

Resource fields#

Field Type Example Description
threads Integer 16 Number of CPU cores
memory String "32G" RAM allocation
gpu Integer 1 Number of GPUs (simple count)
gpu_spec Table See below Detailed GPU specification
disk String "100G" Local disk space — local warning only, never emitted as a scheduler directive
time_limit String "24h" Wall-time limit

Memory and walltime conventions#

One engine-wide parser reads both memory and disk, so what you see in dry-run/run resource output is what the scheduler gets:

  • Bare numbers are megabytes (memory = "4096" = 4 GB) — on every backend. LSF's own default unit is KB, so oxo-flow converts the parsed figure ×1024 and submits an explicit KB number to -M / rusage[mem=…]; the same config cannot schedule 4 GB locally and 4 MB on the cluster.
  • Accepted forms: a bare number, or a single-letter K/M/G/T suffix ("4096", "8G", "16384M", "1T"). Fractional ("16.5G") and multi-letter-suffix ("4096MB", "2GB") figures parse to nothing engine-wide and are rejected at validate time — the same predicate check's E004 diagnostic applies, so a config that loads is a config the pool, the quota, and every scheduler render agree on. Because PBS mem= and SGE h_vmem= default to their own units (words/bytes), those two backends receive the figure with an explicit unit (mem=8192mb, h_vmem=16384M); SLURM's --mem and LSF's converted KB figure need no help (their defaults match or are converted).
  • time_limit accepts the suffix forms 90s, 30m, 24h, 2d and the scheduler-style colon forms 01:00:00, 24:00:00, 1-06:00:00 — one shared parser for the local executor and every backend, so the same value means the same limit everywhere. Bare numbers ("1500") are rejected at validate time: ambiguous between seconds and LSF minutes — write "1500s" or "25m".
  • A present but unparseable walltime is never silently replaced by a default: every backend renders it verbatim (with a warning in the log), so the scheduler rejects the script loudly instead of killing a long job at a hidden 1-hour limit. The local executor logs a warning and falls back to the global timeout.

GPU Specification#

For basic GPU requests, use the gpu field:

[rules.resources]
gpu = 2  # Request 2 GPUs

For advanced GPU configuration, use gpu_spec:

[rules.resources.gpu_spec]
count = 2
model = "a100"       # GPU model (optional — SLURM only)
memory_gb = 40       # Per-GPU memory in GB (optional — SLURM only)

count works on SLURM, PBS, and SGE; model and memory_gb only affect SLURM directives (LSF ignores gpu_spec entirely).

Different schedulers handle GPU requests differently:

Scheduler GPU Directive Notes
SLURM --gres=gpu:2 or --gres=gpu:a100:2:40g Full support for model and memory spec
PBS gpu=2 Basic count only; model selection varies by site
SGE -l gpu=2 Basic count only; requires queue configuration
LSF -gpu "num=2" Basic count only (rendered as a quoted num= run-options string)

SLURM Example#

oxo-flow generates SLURM job scripts automatically. For the align rule above with a sample group batch = ["S1", "S2", "S3"], the script generated for the first instance is:

#!/bin/bash
#SBATCH --chdir=/abs/path/to/workdir
#SBATCH --job-name=align_batch_S1
#SBATCH --cpus-per-task=16
#SBATCH --mem=32G
#SBATCH --time=1-00:00:00
#SBATCH --output=logs/align_batch_S1.out
#SBATCH --error=logs/align_batch_S1.err

set -e

mkdir -p logs
apptainer exec --bind /abs/path/to/workdir:/abs/path/to/workdir docker://biocontainers/bwa:0.7.17 sh -c 'if command -v bash >/dev/null 2>&1; then exec bash -c "$1"; else exec sh -c "$1"; fi' sh 'bwa mem -t 16 ref.fa S1_R1.fastq.gz | samtools sort -o aligned/S1.bam'

Note: one script is generated per rule instance. Wildcards expand before the scripts are written, exactly as they do for run, so a 3-sample scatter produces align_batch_S1.sh, align_batch_S2.sh, and align_batch_S3.sh with concrete paths in each. The instance names match the ones dry-run plans, so a script maps back to a planned rule by name. The sh -c '…' sh '<command>' tail is the bash shim described under Environment Wrapping, and singularity is substituted when apptainer is absent.

Note the workdir pinning at the top of each script (--chdir / cd / -wd): without it SLURM starts a job in the submission directory, PBS in $HOME, and SGE in a queue-configured directory, so a rule's relative output paths would silently land outside the run directory. The examples below show the pin for each backend.


PBS Example#

#!/bin/bash
#PBS -N align_batch_S1
#PBS -l nodes=1:ppn=16,mem=32768mb,walltime=1-00:00:00
#PBS -o logs/align_batch_S1.out
#PBS -e logs/align_batch_S1.err
cd /abs/path/to/workdir || exit 1

set -e

mkdir -p logs
apptainer exec --bind /abs/path/to/workdir:/abs/path/to/workdir docker://biocontainers/bwa:0.7.17 sh -c 'if command -v bash >/dev/null 2>&1; then exec bash -c "$1"; else exec sh -c "$1"; fi' sh 'bwa mem -t 16 ref.fa S1_R1.fastq.gz | samtools sort -o aligned/S1.bam'

SGE Example#

#!/bin/bash
#$ -wd /abs/path/to/workdir
#$ -N align_batch_S1
#$ -pe smp 16
#$ -l h_vmem=32768M
#$ -l h_rt=1-00:00:00
#$ -o logs/align_batch_S1.out
#$ -e logs/align_batch_S1.err

set -e

mkdir -p logs
apptainer exec --bind /abs/path/to/workdir:/abs/path/to/workdir docker://biocontainers/bwa:0.7.17 sh -c 'if command -v bash >/dev/null 2>&1; then exec bash -c "$1"; else exec sh -c "$1"; fi' sh 'bwa mem -t 16 ref.fa S1_R1.fastq.gz | samtools sort -o aligned/S1.bam'

LSF Example#

#!/bin/bash
#BSUB -J align_batch_S1
#BSUB -n 16
#BSUB -R 'rusage[mem=33554432] span[hosts=1]'
#BSUB -M 33554432
#BSUB -W 24:00
#BSUB -o logs/align_batch_S1.out
#BSUB -e logs/align_batch_S1.err

set -e

mkdir -p logs
apptainer exec --bind /abs/path/to/workdir:/abs/path/to/workdir docker://biocontainers/bwa:0.7.17 sh -c 'if command -v bash >/dev/null 2>&1; then exec bash -c "$1"; else exec sh -c "$1"; fi' sh 'bwa mem -t 16 ref.fa S1_R1.fastq.gz | samtools sort -o aligned/S1.bam'

The memory figure is explicit KB ("32G" = 32768 MB × 1024) because LSF's own default unit is KB; span[hosts=1] rides in the -R string so a multi-thread rule stays on one host. -W 24:00 is the parsed form of the rule's time_limit = "24h" — LSF takes [HH:]MM. The workdir pin other backends need is implicit here: an LSF job starts in the directory the script was submitted from.


Environment Wrapping#

When generating cluster scripts, oxo-flow automatically wraps commands through the environment resolver:

Backend Wrapping
Conda conda run --no-capture-output -n <env> bash -c 'export PATH="$CONDA_PREFIX/bin:$PATH"; <command>'
Mamba <mamba\|micromamba\|conda> run -n <env> bash -c 'export PATH="$CONDA_PREFIX/bin:$PATH"; <command>' (--no-capture-output added when the detected binary is conda)
Docker docker run --rm --user $(id -u):$(id -g) -v <workdir>:<workdir> -w <workdir> <image> sh -c '<shim>' sh '<command>'
Singularity / Apptainer <apptainer\|singularity> exec --bind <workdir>:<workdir> <image> sh -c '<shim>' sh '<command>'
Pixi pixi run --manifest-path <pixi.toml> bash -c '<command>'
Venv source <venv>/bin/activate && <command>
Modules module-init block, then module load <mod1> <mod2> && <command>

Three details worth knowing:

  • Container paths are absolute. docker -v and singularity --bind reject relative sources, and a scheduler may start a job from any directory, so oxo-flow renders the resolved workdir. The rule's command is handed to a shim inside the container — if command -v bash >/dev/null 2>&1; then exec bash -c "$1"; else exec sh -c "$1"; fi — because the image's entrypoint shell runs first, and multi-line or pipefail-using scripts need bash, which not every image ships.
  • Paths outside the workdir are bound read-only (Docker backend). A [config] reference genome or index living elsewhere on the shared filesystem is appended to the Docker mount list automatically (-v /data/ref:/data/ref:ro), so the rule can read it from inside the container. The Singularity/Apptainer backend does not add these mounts — bind paths yourself with [cluster.container] bind options if your site needs them.
  • apptainer wins when both it and singularity are installed, so clusters that renamed the binary keep working without configuration.

Environment Examples#

Conda with GPU for deep learning:

[[rules]]
name = "train_model"
input = ["data/train.h5"]
output = ["models/trained.pt"]
environment = { conda = "envs/pytorch.yaml" }
shell = "python train.py --input {input} --output {output}"

[rules.resources]
threads = 8
memory = "64G"
gpu = 2
time_limit = "24h"

The gpu field controls the scheduler allocation only (e.g. #SBATCH --gres=gpu:2) — there is no {resources.gpu} shell placeholder.

Singularity (common on HPC):

[[rules]]
name = "variant_call"
input = ["aligned/{sample}.bam"]
output = ["variants/{sample}.vcf"]
environment = { singularity = "docker://broadinstitute/gatk:4.4.0.0" }
shell = "gatk HaplotypeCaller -I {input} -O {output}"

[rules.resources]
threads = 16
memory = "32G"

Only one backend applies per rule, so combining e.g. singularity with modules in the same rule is rejected outright (the module loads would be silently dropped) — use the module-based example below for module-only setups.

Pixi for reproducible environments:

[[rules]]
name = "qc_check"
input = ["{sample}.fastq.gz"]
output = ["qc/{sample}_fastqc.html"]
environment = { pixi = "envs/pixi.toml" }  # the manifest FILE, not an env name
shell = "fastqc -t {threads} -o qc/ {input}"

[rules.resources]
threads = 4

Pure Module-based (traditional HPC):

[[rules]]
name = "align"
input = ["reads/{sample}.fq"]
output = ["aligned/{sample}.bam"]
environment = { modules = ["bwa/0.7.17", "samtools/1.17", "gcc/11"] }
shell = "bwa mem -t {threads} ref.fa {input} | samtools sort -o {output}"

[rules.resources]
threads = 32
memory = "64G"

Pre-build environments on cluster nodes

Ensure your conda environments, docker images, or singularity containers are available on all cluster nodes before submitting jobs. Use --skip-env-setup when environments are pre-built.


Resource Enforcement#

Local Execution#

When running locally (oxo-flow run), resource constraints are enforced:

  • Check: Before execution, verify resources are available
  • Reserve: Reserve resources before starting the job
  • Release: Release resources after completion (or on failure/timeout)
# Limit to 16 threads and 32GB memory for local execution
oxo-flow run pipeline.oxoflow --max-threads 16 --max-memory 32768

Cluster Execution#

On clusters, the scheduler enforces resources based on the generated directives. oxo-flow does not manage resources during cluster execution — the scheduler handles that.


Best Practices#

Use Singularity on clusters

Most HPC clusters do not allow Docker. Use Singularity instead — oxo-flow handles the conversion automatically when you specify singularity = "docker://...".

Set realistic time limits

Generous wall-time limits prevent premature job termination but may lower scheduling priority. Profile your jobs first.

Use --keep-going for large batches

When running hundreds of samples, use oxo-flow run -k so that a single failure does not abort the entire run.

Aborted in-flight siblings are incomplete, not failed

When a run aborts, rules still in flight are killed and recorded as incomplete (neither completed nor failed) — resume re-runs them from their checkpoint, so no output is silently trusted from a killed job.

Check resource availability

Use sinfo (SLURM), pbsnodes (PBS), or qhost (SGE) to verify available resources before submitting.

Cache environment setup

Use --cache-dir to persist environment setup state across runs for faster startup.


Web UI: Remote Execution over SSH#

The web system can execute runs on a cluster login node while the server itself lives elsewhere (your laptop, a lab server, a container):

  1. Register the connection — Clusters page (or the platform config file's [clusters] section): SSH host, port, user, optional key path, scheduler hint, and a remote_dir for staged runs. Probe verifies connectivity and detects the scheduler (slurm/pbs/sge/lsf) from the remote's installed binaries.
  2. Run with cluster_id — the Run dialog's Execute on cluster selector (or cluster_id in POST /api/runs).
  3. What happens — the workdir is staged to {remote_dir}/runs/{run_id} over a tar stream (no rsync needed); a per-run wrapper launches oxo-flow run remotely under nohup; the server polls the remote .exit-code every 5s; on completion the whole workdir (logs, checkpoint, results) is pulled back, so the web logs, files, preview, and report views work exactly as for local runs. Cancel sends pkill for the per-run wrapper.

Requirements on the remote host: the oxo-flow CLI on PATH (any recent release), tar, and non-interactive SSH access (key auth, BatchMode). Cluster connection management (upsert/delete) is admin-only outside personal mode — in team and hpc mode (they hold shared SSH credentials); every run remains owned by the acting user.

Monitoring Jobs#

After submission, use your cluster's native tools:

# SLURM
squeue -u $USER

# PBS
qstat -u $USER

# SGE
qstat

# LSF
bjobs

Or use oxo-flow's status command with a checkpoint file:

oxo-flow status .oxo-flow/checkpoint.json

oxo-flow cluster status answers two questions. With no ids it lists every job the scheduler currently reports for your user — the natural check right after a submit:

$ oxo-flow cluster status
Cluster: Executing 'squeue -u alice --noheader -o %i|%T'...
Cluster: 2 queued job(s)
  12345: running
  12346: pending

With ids (the ones submit/run printed) it answers exactly those, settling finished jobs from the scheduler's accounting store:

oxo-flow cluster status 12345 12346
oxo-flow cluster cancel 12345

Every cluster action resolves its backend the same way: the --backend / -b flag if given, then $OXO_FLOW_CLUSTER_BACKEND, then slurm. On a SLURM site you never type -b; on a PBS site set the environment variable once (export OXO_FLOW_CLUSTER_BACKEND=pbs). An explicit-but-typoed value is still rejected outright (unknown cluster backend '…' — expected slurm, pbs, sge, or lsf) — the default never swallows a bad name.


Driver Defaults and the [cluster] Profile Block#

run --profile submits and tracks jobs through the backend driver, whose built-in defaults are:

Setting Default Meaning
max_submitted 50 Jobs in flight at once (pending + running); submissions top up to this cap as slots free
max_array_size 1001 The scheduler's array limit; larger scatter groups are chunked into several arrays
poll_interval 5s Delay between scheduler polls

Override them in the profile's [cluster] block:

[cluster]
backend       = "slurm"   # required: slurm | pbs | sge | lsf
partition     = "compute"
account       = "lab01"
walltime      = "24h"     # a rule's own time_limit wins
max_submitted = 100
max_array_size = 1001
poll_interval = "30s"     # duration string
extra_args    = ["--exclusive", "--constraint=haswell"]  # verbatim scheduler args

extra_args entries are emitted as scheduler directives verbatim and unvalidated — a typo reaches the scheduler as written (the same applies to cluster submit --extra-arg). Cluster CLI actions resolve --backend as flag > $OXO_FLOW_CLUSTER_BACKEND > slurm (see Monitoring Jobs); the profile block's backend key remains required and is validated the same strict way.


See Also#

Scheduling Fairness#

Cluster submission honors the same fairness contract as local execution:

  • Priority is honored: each round, ready rules are submitted in effective-priority order — declared priority plus priority aging (every round a ready rule spends beyond the max_submitted cap gains +1). A producer parked at the cap therefore outranks fresh high-priority rules after priority gap rounds — priority-inversion starvation cannot persist (issue #134).
  • In-flight jobs are cancelled on failure: when a required rule fails, the driver cancels every job already submitted (scancel/qdel/bkill) before the run exits — no orphaned cluster jobs.
  • Waits inside the scheduler queue are the scheduler's own ordering and are visible via squeue/qstat/bjobs (job ids are recorded per rule under the run directory). The local 60-second "waiting for resources:" diagnostics have no cluster counterpart by design — the pool lives on the scheduler, not in the engine.