Run on a Cluster#
This guide explains how to execute oxo-flow workflows on HPC clusters using SLURM, PBS, SGE, and LSF backends.
Overview#
oxo-flow's cluster module translates each rule into a cluster job submission. Resource requirements declared in the .oxoflow file (threads, memory, gpu, time_limit) are mapped to the appropriate scheduler directives. The disk field is not mapped to any scheduler directive — it only produces a local warning during oxo-flow run.
There are two ways to reach a scheduler, and they suit different jobs:
| What it does | Use when | |
|---|---|---|
run --profile <NAME> |
Submits, tracks jobs to completion, and updates the checkpoint | Normal execution — you want the workflow to run |
cluster submit |
Writes job scripts and stops | You want to inspect or hand-edit scripts before submitting |
run --profile is the everyday path; it inherits everything run does — wildcard expansion, checkpoint/resume, --samples, --rerun, config-change invalidation. See Cluster submission for the [cluster] profile block. cluster submit remains the escape hatch and is documented below.
Environment wrapping is applied automatically — conda, mamba, docker, singularity/apptainer, pixi, venv, and modules environments are wrapped in the generated scripts. Each rule applies one backend (the first declared in the resolver order mamba → conda → pixi → docker → singularity → venv → modules), so declaring e.g. both singularity and modules fails loudly instead of silently dropping the module loads.
Supported Schedulers#
| Scheduler | Status | Directive prefix |
|---|---|---|
| SLURM | Supported | #SBATCH |
| PBS/Torque | Supported | #PBS |
| SGE | Supported | #$ |
| LSF | Supported | #BSUB |
Declaring Resources#
Set resource requirements per rule:
[[rules]]
name = "align"
input = ["{sample}_R1.fastq.gz"]
output = ["aligned/{sample}.bam"]
environment = { singularity = "docker://biocontainers/bwa:0.7.17" }
shell = "bwa mem -t {threads} ref.fa {input} | samtools sort -o {output}"
[rules.resources]
threads = 16
memory = "32G"
time_limit = "24h"
Note on
gpu:gpu = 0means "no GPUs" — the engine omits the GPU directive entirely, exactly as if the field were absent. Only declare it when you need at least one GPU (gpu = 1or higher).
Resource fields#
| Field | Type | Example | Description |
|---|---|---|---|
threads |
Integer | 16 |
Number of CPU cores |
memory |
String | "32G" |
RAM allocation |
gpu |
Integer | 1 |
Number of GPUs (simple count) |
gpu_spec |
Table | See below | Detailed GPU specification |
disk |
String | "100G" |
Local disk space — local warning only, never emitted as a scheduler directive |
time_limit |
String | "24h" |
Wall-time limit |
Memory and walltime conventions#
One engine-wide parser reads both memory and disk, so what you see in
dry-run/run resource output is what the scheduler gets:
- Bare numbers are megabytes (
memory = "4096"= 4 GB) — on every backend. LSF's own default unit is KB, so oxo-flow converts the parsed figure ×1024 and submits an explicit KB number to-M/rusage[mem=…]; the same config cannot schedule 4 GB locally and 4 MB on the cluster. - Accepted forms: a bare number, or a single-letter
K/M/G/Tsuffix ("4096","8G","16384M","1T"). Fractional ("16.5G") and multi-letter-suffix ("4096MB","2GB") figures parse to nothing engine-wide and are rejected at validate time — the same predicatecheck's E004 diagnostic applies, so a config that loads is a config the pool, the quota, and every scheduler render agree on. Because PBSmem=and SGEh_vmem=default to their own units (words/bytes), those two backends receive the figure with an explicit unit (mem=8192mb,h_vmem=16384M); SLURM's--memand LSF's converted KB figure need no help (their defaults match or are converted). time_limitaccepts the suffix forms90s,30m,24h,2dand the scheduler-style colon forms01:00:00,24:00:00,1-06:00:00— one shared parser for the local executor and every backend, so the same value means the same limit everywhere. Bare numbers ("1500") are rejected at validate time: ambiguous between seconds and LSF minutes — write"1500s"or"25m".- A present but unparseable walltime is never silently replaced by a default: every backend renders it verbatim (with a warning in the log), so the scheduler rejects the script loudly instead of killing a long job at a hidden 1-hour limit. The local executor logs a warning and falls back to the global timeout.
GPU Specification#
For basic GPU requests, use the gpu field:
For advanced GPU configuration, use gpu_spec:
[rules.resources.gpu_spec]
count = 2
model = "a100" # GPU model (optional — SLURM only)
memory_gb = 40 # Per-GPU memory in GB (optional — SLURM only)
count works on SLURM, PBS, and SGE; model and memory_gb only affect
SLURM directives (LSF ignores gpu_spec entirely).
Different schedulers handle GPU requests differently:
| Scheduler | GPU Directive | Notes |
|---|---|---|
| SLURM | --gres=gpu:2 or --gres=gpu:a100:2:40g |
Full support for model and memory spec |
| PBS | gpu=2 |
Basic count only; model selection varies by site |
| SGE | -l gpu=2 |
Basic count only; requires queue configuration |
| LSF | -gpu "num=2" |
Basic count only (rendered as a quoted num= run-options string) |
SLURM Example#
oxo-flow generates SLURM job scripts automatically. For the align rule
above with a sample group batch = ["S1", "S2", "S3"], the script generated
for the first instance is:
#!/bin/bash
#SBATCH --chdir=/abs/path/to/workdir
#SBATCH --job-name=align_batch_S1
#SBATCH --cpus-per-task=16
#SBATCH --mem=32G
#SBATCH --time=1-00:00:00
#SBATCH --output=logs/align_batch_S1.out
#SBATCH --error=logs/align_batch_S1.err
set -e
mkdir -p logs
apptainer exec --bind /abs/path/to/workdir:/abs/path/to/workdir docker://biocontainers/bwa:0.7.17 sh -c 'if command -v bash >/dev/null 2>&1; then exec bash -c "$1"; else exec sh -c "$1"; fi' sh 'bwa mem -t 16 ref.fa S1_R1.fastq.gz | samtools sort -o aligned/S1.bam'
Note: one script is generated per rule instance. Wildcards expand before
the scripts are written, exactly as they do for run, so a 3-sample scatter
produces align_batch_S1.sh, align_batch_S2.sh, and align_batch_S3.sh
with concrete paths in each. The instance names match the ones dry-run
plans, so a script maps back to a planned rule by name. The sh -c '…' sh
'<command>' tail is the bash shim described under
Environment Wrapping, and singularity is
substituted when apptainer is absent.
Note the workdir pinning at the top of each script (--chdir / cd /
-wd): without it SLURM starts a job in the submission directory, PBS in
$HOME, and SGE in a queue-configured directory, so a rule's relative
output paths would silently land outside the run directory. The examples
below show the pin for each backend.
PBS Example#
#!/bin/bash
#PBS -N align_batch_S1
#PBS -l nodes=1:ppn=16,mem=32768mb,walltime=1-00:00:00
#PBS -o logs/align_batch_S1.out
#PBS -e logs/align_batch_S1.err
cd /abs/path/to/workdir || exit 1
set -e
mkdir -p logs
apptainer exec --bind /abs/path/to/workdir:/abs/path/to/workdir docker://biocontainers/bwa:0.7.17 sh -c 'if command -v bash >/dev/null 2>&1; then exec bash -c "$1"; else exec sh -c "$1"; fi' sh 'bwa mem -t 16 ref.fa S1_R1.fastq.gz | samtools sort -o aligned/S1.bam'
SGE Example#
#!/bin/bash
#$ -wd /abs/path/to/workdir
#$ -N align_batch_S1
#$ -pe smp 16
#$ -l h_vmem=32768M
#$ -l h_rt=1-00:00:00
#$ -o logs/align_batch_S1.out
#$ -e logs/align_batch_S1.err
set -e
mkdir -p logs
apptainer exec --bind /abs/path/to/workdir:/abs/path/to/workdir docker://biocontainers/bwa:0.7.17 sh -c 'if command -v bash >/dev/null 2>&1; then exec bash -c "$1"; else exec sh -c "$1"; fi' sh 'bwa mem -t 16 ref.fa S1_R1.fastq.gz | samtools sort -o aligned/S1.bam'
LSF Example#
#!/bin/bash
#BSUB -J align_batch_S1
#BSUB -n 16
#BSUB -R 'rusage[mem=33554432] span[hosts=1]'
#BSUB -M 33554432
#BSUB -W 24:00
#BSUB -o logs/align_batch_S1.out
#BSUB -e logs/align_batch_S1.err
set -e
mkdir -p logs
apptainer exec --bind /abs/path/to/workdir:/abs/path/to/workdir docker://biocontainers/bwa:0.7.17 sh -c 'if command -v bash >/dev/null 2>&1; then exec bash -c "$1"; else exec sh -c "$1"; fi' sh 'bwa mem -t 16 ref.fa S1_R1.fastq.gz | samtools sort -o aligned/S1.bam'
The memory figure is explicit KB ("32G" = 32768 MB × 1024) because LSF's
own default unit is KB; span[hosts=1] rides in the -R string so a
multi-thread rule stays on one host. -W 24:00 is the parsed form of the
rule's time_limit = "24h" — LSF takes [HH:]MM. The workdir pin other
backends need is implicit here: an LSF job starts in the directory the
script was submitted from.
Environment Wrapping#
When generating cluster scripts, oxo-flow automatically wraps commands through the environment resolver:
| Backend | Wrapping |
|---|---|
| Conda | conda run --no-capture-output -n <env> bash -c 'export PATH="$CONDA_PREFIX/bin:$PATH"; <command>' |
| Mamba | <mamba\|micromamba\|conda> run -n <env> bash -c 'export PATH="$CONDA_PREFIX/bin:$PATH"; <command>' (--no-capture-output added when the detected binary is conda) |
| Docker | docker run --rm --user $(id -u):$(id -g) -v <workdir>:<workdir> -w <workdir> <image> sh -c '<shim>' sh '<command>' |
| Singularity / Apptainer | <apptainer\|singularity> exec --bind <workdir>:<workdir> <image> sh -c '<shim>' sh '<command>' |
| Pixi | pixi run --manifest-path <pixi.toml> bash -c '<command>' |
| Venv | source <venv>/bin/activate && <command> |
| Modules | module-init block, then module load <mod1> <mod2> && <command> |
Three details worth knowing:
- Container paths are absolute.
docker -vandsingularity --bindreject relative sources, and a scheduler may start a job from any directory, so oxo-flow renders the resolved workdir. The rule's command is handed to a shim inside the container —if command -v bash >/dev/null 2>&1; then exec bash -c "$1"; else exec sh -c "$1"; fi— because the image's entrypoint shell runs first, and multi-line orpipefail-using scripts need bash, which not every image ships. - Paths outside the workdir are bound read-only (Docker backend). A
[config]reference genome or index living elsewhere on the shared filesystem is appended to the Docker mount list automatically (-v /data/ref:/data/ref:ro), so the rule can read it from inside the container. The Singularity/Apptainer backend does not add these mounts — bind paths yourself with[cluster.container]bind options if your site needs them. apptainerwins when both it andsingularityare installed, so clusters that renamed the binary keep working without configuration.
Environment Examples#
Conda with GPU for deep learning:
[[rules]]
name = "train_model"
input = ["data/train.h5"]
output = ["models/trained.pt"]
environment = { conda = "envs/pytorch.yaml" }
shell = "python train.py --input {input} --output {output}"
[rules.resources]
threads = 8
memory = "64G"
gpu = 2
time_limit = "24h"
The gpu field controls the scheduler allocation only (e.g.
#SBATCH --gres=gpu:2) — there is no {resources.gpu} shell placeholder.
Singularity (common on HPC):
[[rules]]
name = "variant_call"
input = ["aligned/{sample}.bam"]
output = ["variants/{sample}.vcf"]
environment = { singularity = "docker://broadinstitute/gatk:4.4.0.0" }
shell = "gatk HaplotypeCaller -I {input} -O {output}"
[rules.resources]
threads = 16
memory = "32G"
Only one backend applies per rule, so combining e.g. singularity with
modules in the same rule is rejected outright (the module loads would be
silently dropped) — use the module-based example below for module-only
setups.
Pixi for reproducible environments:
[[rules]]
name = "qc_check"
input = ["{sample}.fastq.gz"]
output = ["qc/{sample}_fastqc.html"]
environment = { pixi = "envs/pixi.toml" } # the manifest FILE, not an env name
shell = "fastqc -t {threads} -o qc/ {input}"
[rules.resources]
threads = 4
Pure Module-based (traditional HPC):
[[rules]]
name = "align"
input = ["reads/{sample}.fq"]
output = ["aligned/{sample}.bam"]
environment = { modules = ["bwa/0.7.17", "samtools/1.17", "gcc/11"] }
shell = "bwa mem -t {threads} ref.fa {input} | samtools sort -o {output}"
[rules.resources]
threads = 32
memory = "64G"
Pre-build environments on cluster nodes
Ensure your conda environments, docker images, or singularity containers are available on all cluster nodes before submitting jobs. Use --skip-env-setup when environments are pre-built.
Resource Enforcement#
Local Execution#
When running locally (oxo-flow run), resource constraints are enforced:
- Check: Before execution, verify resources are available
- Reserve: Reserve resources before starting the job
- Release: Release resources after completion (or on failure/timeout)
# Limit to 16 threads and 32GB memory for local execution
oxo-flow run pipeline.oxoflow --max-threads 16 --max-memory 32768
Cluster Execution#
On clusters, the scheduler enforces resources based on the generated directives. oxo-flow does not manage resources during cluster execution — the scheduler handles that.
Best Practices#
Use Singularity on clusters
Most HPC clusters do not allow Docker. Use Singularity instead — oxo-flow handles the conversion automatically when you specify singularity = "docker://...".
Set realistic time limits
Generous wall-time limits prevent premature job termination but may lower scheduling priority. Profile your jobs first.
Use --keep-going for large batches
When running hundreds of samples, use oxo-flow run -k so that a single failure does not abort the entire run.
Aborted in-flight siblings are incomplete, not failed
When a run aborts, rules still in flight are killed and recorded as
incomplete (neither completed nor failed) — resume re-runs them
from their checkpoint, so no output is silently trusted from a killed
job.
Check resource availability
Use sinfo (SLURM), pbsnodes (PBS), or qhost (SGE) to verify available resources before submitting.
Cache environment setup
Use --cache-dir to persist environment setup state across runs for faster startup.
Web UI: Remote Execution over SSH#
The web system can execute runs on a cluster login node while the server itself lives elsewhere (your laptop, a lab server, a container):
- Register the connection — Clusters page (or the platform config
file's
[clusters]section): SSH host, port, user, optional key path, scheduler hint, and aremote_dirfor staged runs. Probe verifies connectivity and detects the scheduler (slurm/pbs/sge/lsf) from the remote's installed binaries. - Run with
cluster_id— the Run dialog's Execute on cluster selector (orcluster_idinPOST /api/runs). - What happens — the workdir is staged to
{remote_dir}/runs/{run_id}over a tar stream (no rsync needed); a per-run wrapper launchesoxo-flow runremotely under nohup; the server polls the remote.exit-codeevery 5s; on completion the whole workdir (logs, checkpoint, results) is pulled back, so the web logs, files, preview, and report views work exactly as for local runs. Cancel sendspkillfor the per-run wrapper.
Requirements on the remote host: the oxo-flow CLI on PATH (any recent
release), tar, and non-interactive SSH access (key auth, BatchMode).
Cluster connection management (upsert/delete) is admin-only outside
personal mode — in team and hpc mode (they hold shared SSH credentials);
every run remains owned by the acting user.
Monitoring Jobs#
After submission, use your cluster's native tools:
Or use oxo-flow's status command with a checkpoint file:
oxo-flow cluster status answers two questions. With no ids it lists
every job the scheduler currently reports for your user — the natural check
right after a submit:
$ oxo-flow cluster status
Cluster: Executing 'squeue -u alice --noheader -o %i|%T'...
Cluster: 2 queued job(s)
12345: running
12346: pending
With ids (the ones submit/run printed) it answers exactly those, settling finished jobs from the scheduler's accounting store:
Every cluster action resolves its backend the same way: the --backend /
-b flag if given, then $OXO_FLOW_CLUSTER_BACKEND, then slurm. On a
SLURM site you never type -b; on a PBS site set the environment variable
once (export OXO_FLOW_CLUSTER_BACKEND=pbs). An explicit-but-typoed value
is still rejected outright (unknown cluster backend '…' — expected slurm,
pbs, sge, or lsf) — the default never swallows a bad name.
Driver Defaults and the [cluster] Profile Block#
run --profile submits and tracks jobs through the backend driver, whose
built-in defaults are:
| Setting | Default | Meaning |
|---|---|---|
max_submitted |
50 |
Jobs in flight at once (pending + running); submissions top up to this cap as slots free |
max_array_size |
1001 |
The scheduler's array limit; larger scatter groups are chunked into several arrays |
poll_interval |
5s |
Delay between scheduler polls |
Override them in the profile's [cluster] block:
[cluster]
backend = "slurm" # required: slurm | pbs | sge | lsf
partition = "compute"
account = "lab01"
walltime = "24h" # a rule's own time_limit wins
max_submitted = 100
max_array_size = 1001
poll_interval = "30s" # duration string
extra_args = ["--exclusive", "--constraint=haswell"] # verbatim scheduler args
extra_args entries are emitted as scheduler directives verbatim and
unvalidated — a typo reaches the scheduler as written (the same applies to
cluster submit --extra-arg). Cluster CLI actions resolve --backend as
flag > $OXO_FLOW_CLUSTER_BACKEND > slurm (see
Monitoring Jobs); the profile block's backend key
remains required and is validated the same strict way.
See Also#
- Architecture: Cluster backends — internal cluster module design
- Environment System — Singularity and Docker on HPC
runcommand —--max-threads,--max-memory,--skip-env-setup,--cache-dirclustercommand — cluster submission reference
Scheduling Fairness#
Cluster submission honors the same fairness contract as local execution:
- Priority is honored: each round, ready rules are submitted in
effective-priority order — declared
priorityplus priority aging (every round a ready rule spends beyond themax_submittedcap gains +1). A producer parked at the cap therefore outranks fresh high-priority rules afterpriority gaprounds — priority-inversion starvation cannot persist (issue #134). - In-flight jobs are cancelled on failure: when a required rule
fails, the driver cancels every job already submitted
(
scancel/qdel/bkill) before the run exits — no orphaned cluster jobs. - Waits inside the scheduler queue are the scheduler's own
ordering and are visible via
squeue/qstat/bjobs(job ids are recorded per rule under the run directory). The local 60-second "waiting for resources:" diagnostics have no cluster counterpart by design — the pool lives on the scheduler, not in the engine.