oxo-flow cluster#
Manage cluster job submission and monitoring.
Everyday path:
run --profile <NAME>submits to a scheduler and tracks jobs to completion when the profile carries a[cluster]block — see Cluster submission. The commands below remain the manual escape hatch: inspect scripts before submitting, or cancel/collect jobs after an interrupted run.
Usage#
Actions#
| Action | Description |
|---|---|
submit |
Submit a workflow to a cluster scheduler |
status |
Show the status of submitted cluster jobs |
cancel |
Cancel submitted cluster jobs |
logs |
Fetch the scheduler's job record (SLURM: sacct accounting output) for a submitted cluster job |
Arguments#
| Argument | Description |
|---|---|
<WORKFLOW> |
Path to the .oxoflow workflow file (for submit) |
[JOB_IDS]... |
Cluster job IDs from a submit/run output — for status: omitted, the command lists your queued jobs; given, it answers exactly those ids. Optional for cancel |
<JOB_ID> |
Job ID (for logs) |
Options#
Every action accepts --backend / -b:
| Option | Short | Default | Description |
|---|---|---|---|
--backend |
-b |
$OXO_FLOW_CLUSTER_BACKEND, then slurm |
Cluster backend (slurm, pbs, sge, lsf) — on a SLURM site you never need to type it |
An explicit value is always validated — -b slrm fails with
unknown cluster backend 'slrm' — expected slurm, pbs, sge, or lsf rather
than silently querying SLURM. The default applies only when the flag is
omitted entirely. On a non-SLURM site set the default once in your shell
profile:
Options (Submit)#
| Option | Short | Default | Description |
|---|---|---|---|
--queue |
-q |
— | Partition / queue name |
--account |
-a |
— | Account / project name |
--walltime |
— | — | Wall-time limit for every job (24h, 2d, or 24:00:00) |
--extra-arg |
— | — | Extra scheduler argument, passed through verbatim (repeatable) |
--output |
-o |
cluster_scripts |
Directory for generated scripts |
--target |
-t |
— | Target rule(s) to execute |
--module |
— | — | Run one include module plus the producers of its declared inputs (repeatable; unions with --target). Module names are the include's name field or its file stem |
--with-dependencies |
— | — | Generate dependency-aware submit script with job chains |
--dry-run |
— | — | Generate and write the scripts but submit nothing (and skip submit.sh even with --with-dependencies) |
One script is written per rule instance: wildcards expand first, so a
scatter rule over three samples yields three scripts whose names match the
instances dry-run plans.
A rule's own time_limit beats --walltime. --extra-arg values are
emitted as scheduler directives verbatim and are not validated — a typo
reaches the scheduler as written.
Examples#
Submit to SLURM#
Submit to PBS/Torque#
Submit to SGE (Sun Grid Engine)#
Submit to LSF#
Submit with queue and account#
Submit with environment support#
# If your workflow uses conda environments, the generated scripts
# will automatically include conda activation commands
oxo-flow cluster submit pipeline.oxoflow -b slurm -q compute
Submit with a wall-time limit and site-specific flags#
oxo-flow cluster submit pipeline.oxoflow -b slurm -q compute \
--walltime 24h --extra-arg --exclusive --extra-arg --constraint=haswell
Submit with job dependencies#
# Generate scripts with automatic dependency chain setup
# Creates a submit.sh wrapper script that handles job submission order
oxo-flow cluster submit pipeline.oxoflow -b slurm -q compute --with-dependencies
# Submit the generated wrapper script
bash cluster_scripts/submit.sh
The wrapper submits through an oxo_submit helper that captures the bare
scheduler job id — sbatch --parsable on SLURM, sentence parsing on SGE and
LSF — before chaining it into the next job's dependency flag. Dependencies
are wired per instance, so sample 2's stats waits on sample 2's align
rather than on every sample's.
Submit specific target rules#
# Only generate scripts for specific rules and their dependencies
oxo-flow cluster submit pipeline.oxoflow -b slurm -q compute -t align -t call_variants
Dry run mode#
# Write the scripts without submitting anything
oxo-flow cluster submit pipeline.oxoflow -b slurm -q compute --dry-run
--dry-run still writes every script to the output directory — it prints
(dry-run) generating … job scripts … nothing is submitted and ends with the
submission line to review, e.g. sbatch cluster_scripts/*.sh.
Check job status#
With no ids, status lists your queued jobs (all jobs the scheduler
reports for your user) — the natural first question after a submit:
$ oxo-flow cluster status
Cluster: Executing 'squeue -u alice --noheader -o %i|%T'...
Cluster: 2 queued job(s)
12345: running
12346: pending
Pass the ids the submit/run step printed to answer exactly those — including finished ones, which are settled from the scheduler's accounting store:
Cancel specific jobs#
Fetch a job's scheduler record#
SLURM prints the job's sacct record
(JobID,State,ExitCode,Elapsed,MaxRSS,TotalCPU); PBS/SGE/LSF are
best-effort (qstat -f / qacct / bacct). Requires the
scheduler's client commands on PATH.
cluster logs on SLURM reads sacct directly and does NOT fall back to
scontrol: on clusters without slurmdbd (accounting storage disabled),
sacct exits non-zero and this command errors with
'sacct' exited 1: Slurm accounting storage is disabled (live-verified on
such a cluster). The scontrol show job <id> fallback mentioned below
belongs to run-time settlement inside oxo-flow run, not to this
command.
On SLURM clusters without slurmdbd, settlement in oxo-flow run falls
back to scontrol show job <id>, which reads the controller's in-memory
record (state + exit code; no Elapsed/RSS/CPU). A job that has left both the
live queue and every probe for longer than the blind-settlement window
settles as failed with an unknown exit code and a warning naming the
rule — the run ends instead of polling forever, and the operator verifies
via the rule's output files.
Output#
Basic Output#
oxo-flow v0.23.2 — Rust-native bioinformatics pipeline engine
Cluster: Generating slurm job scripts for 5 rule instances
✓ cluster_scripts/fastqc.sh
✓ cluster_scripts/trim_reads.sh
✓ cluster_scripts/bwa_align_S1.sh
✓ cluster_scripts/bwa_align_S2.sh
✓ cluster_scripts/bwa_align_S3.sh
Done: 5 scripts written to cluster_scripts
Submit with: sbatch cluster_scripts/*.sh
With Dependencies Output#
oxo-flow v0.23.2 — Rust-native bioinformatics pipeline engine
Cluster: Generating slurm job scripts for 5 rule instances
✓ cluster_scripts/fastqc.sh
✓ cluster_scripts/trim_reads.sh
✓ cluster_scripts/bwa_align_S1.sh
✓ cluster_scripts/bwa_align_S2.sh
✓ cluster_scripts/bwa_align_S3.sh
✓ cluster_scripts/submit.sh (dependency-aware submit script)
Done: 6 scripts written to cluster_scripts
Submit with: bash cluster_scripts/submit.sh
Generated Script Example#
For a workflow rule with conda environment, different backends produce different scripts
(threads/memory come from the rule's resources; the -q/-a flags add queue/account
directives; --walltime — or a rule's time_limit, which wins — adds a time directive,
#SBATCH --time= on SLURM and walltime= on PBS):
SLURM Script#
#!/bin/bash
#SBATCH --job-name=bwa_align
#SBATCH --cpus-per-task=16
#SBATCH --mem=32G
#SBATCH --partition=compute
#SBATCH --output=logs/bwa_align.out
#SBATCH --error=logs/bwa_align.err
set -e
mkdir -p logs
conda run --no-capture-output -n bwa_env bash -c 'export PATH="$CONDA_PREFIX/bin:$PATH"; bwa mem -t 16 ref.fa reads.fq > aligned.sam'
PBS/Torque Script#
#!/bin/bash
#PBS -N bwa_align
#PBS -l nodes=1:ppn=16,mem=32G
#PBS -o logs/bwa_align.out
#PBS -e logs/bwa_align.err
set -e
mkdir -p logs
conda run --no-capture-output -n bwa_env bash -c 'export PATH="$CONDA_PREFIX/bin:$PATH"; bwa mem -t 16 ref.fa reads.fq > aligned.sam'
SGE Script#
#!/bin/bash
#$ -N bwa_align
#$ -pe smp 16
#$ -l h_vmem=32G
#$ -o logs/bwa_align.out
#$ -e logs/bwa_align.err
set -e
mkdir -p logs
conda run --no-capture-output -n bwa_env bash -c 'export PATH="$CONDA_PREFIX/bin:$PATH"; bwa mem -t 16 ref.fa reads.fq > aligned.sam'
LSF Script#
#!/bin/bash
#BSUB -J bwa_align
#BSUB -n 16
#BSUB -R 'rusage[mem=33554432] span[hosts=1]'
#BSUB -M 33554432
#BSUB -o logs/bwa_align.out
#BSUB -e logs/bwa_align.err
set -e
mkdir -p logs
conda run --no-capture-output -n bwa_env bash -c 'export PATH="$CONDA_PREFIX/bin:$PATH"; bwa mem -t 16 ref.fa reads.fq > aligned.sam'
Memory is emitted as an explicit KB figure (32G → 33554432) in both
-R 'rusage[mem=…]' and -M: LSF's documented mem_spec unit set is
KB/MB/GB with KB as the default, and the bare 32G form oxo-flow stores
internally is not a valid mem_spec. The rusage is what actually places
the job on a host with room; multi-threaded rules also carry
span[hosts=1] inside the same -R string.
Environment wrapping is applied automatically: conda rules are wrapped in
conda run --no-capture-output -n <env> bash -c 'export PATH="$CONDA_PREFIX/bin:$PATH"; ...', docker rules in
docker run --rm --user $(id -u):$(id -g) ... <image> sh -c '<bash shim>' sh '...', singularity/apptainer rules in
<apptainer|singularity> exec --bind <workdir>:<workdir> ... <image> sh -c '<bash shim>' sh '...', and rules
without an environment run the command directly. The shim re-execs the
command under bash when the image ships it — see
Environment Wrapping.
Dependency-Aware Submit Script#
When using --with-dependencies, oxo-flow generates a submit.sh wrapper that handles job submission order:
#!/bin/bash
# Auto-generated dependency-aware submit script
# Generated by oxo-flow
set -e
# Track job IDs
declare -A JOB_IDS
# Submit one script and echo its scheduler job id.
oxo_submit() {
local out id
out=$(sbatch --parsable "$@") || return $?
id=${out%%;*}
if [ -z "$id" ]; then
echo "oxo-flow: cannot parse job id from: $out" >&2
return 1
fi
printf '%s' "$id"
}
echo 'Submitting fastqc...'
JOB_IDS[fastqc]=$(oxo_submit cluster_scripts/fastqc.sh)
echo " Submitted fastqc as job ID: ${JOB_IDS[fastqc]}"
echo 'Submitting trim_reads...'
JOB_IDS[trim_reads]=$(oxo_submit --dependency=afterok:${JOB_IDS[fastqc]} cluster_scripts/trim_reads.sh)
echo " Submitted trim_reads as job ID: ${JOB_IDS[trim_reads]}"
echo 'Submitting bwa_align...'
JOB_IDS[bwa_align]=$(oxo_submit --dependency=afterok:${JOB_IDS[trim_reads]} cluster_scripts/bwa_align.sh)
echo " Submitted bwa_align as job ID: ${JOB_IDS[bwa_align]}"
echo 'All jobs submitted successfully!'
echo 'Job ID mapping:'
for name in "${!JOB_IDS[@]}"; do
echo " $name: ${JOB_IDS[$name]}"
done
Different backends use different dependency syntax:
| Backend | Dependency Flag |
|---|---|
| SLURM | --dependency=afterok:jobid |
| PBS | -W depend=afterok:jobid |
| SGE | -hold_jid jobid |
| LSF | -w "ended(jobid)" |
Element-wise array chaining (SLURM)#
When a downstream rule has exactly one upstream dependency and BOTH scripts
declare the same straight #SBATCH --array=START-END range (typically via
--extra-arg "--array=1-N"), the wrapper chains them element-wise:
JOB_IDS[trim_reads]=$(oxo_submit --dependency=aftercorr:${JOB_IDS[fastqc]} cluster_scripts/trim_reads.sh)
aftercorr pairs array elements by index, so element 1 of the downstream
starts as soon as element 1 of the upstream finishes instead of waiting for
the whole array — a pipelining win on straight scatter chains (measured
1.33× makespan improvement on a real cluster). Any other shape keeps the
ready-batch afterok dependency: fan-in (multiple dependencies), mismatched
index ranges, scalar↔array mixes, or a non-array upstream. A %throttle
suffix does not block the match — it bounds concurrency, not the index set.
Notes#
submitgenerates shell scripts tailored for the specified cluster backend- Script generation goes through the
ExecutorBackendtrait (Execution Backends) — the same render layer the live submission path uses, so generated scripts and submitted scripts can never drift apart - The
logs/directory referenced by the scripts'--outputdirectives is created at submit time (both byoxo-flow run --profileand by the engine's live submission path): schedulers open the script's--outputfile at job launch, before the script body runs, so a missing directory fails the job instantly. When you submit generated scripts yourself (cluster submit --dry-runoutput), create the directory first. - Resource requirements (threads, memory, gpu) from the workflow are automatically translated to cluster directives
- Environment wrapping is applied automatically — conda, docker, singularity, pixi, venv, and module environments are properly wrapped in the generated scripts
status,cancel, andlogsactively execute native cluster commands (squeue,scancel, …).statuscaptures the scheduler's output and parses it into one normalized state per job (pending / running / completed / failed / cancelled / unknown); for ids you request explicitly, jobs no longer in the queue fall back to the scheduler's accounting store (sacct,qstat -x,qacct,bacct) so finished jobs report their real final state instead of "not in the queue". The no-ids listing shows only live (queued or running) jobs — that is what the scheduler itself reports.- Ensure the required environments (conda envs, docker images, etc.) are available on cluster nodes before submitting
- Use
--with-dependenciesfor workflows where rules depend on each other — this ensures proper execution order