oxo-flow run#
Execute a workflow.
Usage#
Arguments#
| Argument | Description |
|---|---|
[WORKFLOW] |
Path to the .oxoflow workflow file. Optional — if not specified, auto-discovery searches for: (1) main.oxoflow in current directory, (2) alphabetically first *.oxoflow file in current directory. |
[KEY=VALUE]... |
Direct config overrides: KEY=VALUE and --KEY=VALUE work for any key; the --KEY VALUE space form also works for any key present in [config] (plain key = "value" entries included) — a key absent from [config] is a hard error (an unknown flag would otherwise be indistinguishable from a mistyped option; a typo'd --key=val after the workflow still sets a config var, so prefer the = forms for undeclared keys). An override key that differs from a declared key only by hyphens vs underscores is canonicalized onto the declared spelling — --api-token=x sets the declared key api_token (and inherits its sensitive masking); keys matching nothing declared keep the literal spelling |
Options#
| Option | Short | Default | Description |
|---|---|---|---|
--jobs |
-j |
1 |
Maximum number of concurrent jobs |
--keep-going |
-k |
— | Continue execution when a job fails |
--json |
— | — | Emit a machine-readable JSON summary on stdout after the run (command, status, results counts, resources). The document is emitted on every exit path — completed, failed, and aborted (preflight failures, budget breaches, cluster runs, plain-failure aborts) — so --json consumers never get zero bytes. status is "completed" only when the run finished with zero required failures; every other exit is "failed". Human progress and result output on stderr is not suppressed — this flag appends the summary, it does not switch the run to machine-only mode. resources lists every benchmarked rule (rule, status, wall_time_secs, peak_rss_mb, cpu_seconds, retries) — the same data as the report's Benchmarks table; it is empty on aborted paths and on cluster runs (use the report's sacct-based Resource Accounting section there) |
--workdir |
-d |
Workflow file's directory | Working directory for execution |
--log-file |
— | .oxo-flow/logs/oxo-flow.log in the workdir |
Write the run log to a custom path instead (relative paths resolve against the workdir; previous logs rotate to PATH.1 … PATH.9) |
--target |
-t |
All rules | Run only specific target rules |
--module |
— | — | Run one include module and the producers of its declared inputs (repeatable; unions with --target). Module names are the include's name field or its file stem (see Partial module runs) |
--retry |
-r |
0 |
Number of times to retry failed jobs |
--timeout |
— | 0 (disabled) |
Timeout per job in seconds, or a duration like 1h/30m |
--max-threads |
— | 0 (auto-detect) |
Maximum CPU threads available for execution |
--max-memory |
— | 0 (auto-detect) |
Maximum memory in MB available for execution |
--skip-env-setup |
— | — | Skip environment setup (assume environments are ready) |
--skip-ref-build |
— | — | Skip automatic reference/index building (assume pre-built) |
--cache-dir |
— | — | Directory for caching environment setup state (entries untouched for 90 days are cleaned up after each run; override with the cache_max_age_days config key, 0 disables aging) |
--resume-failed |
— | — | Resume only failed rules from a previous run. Also clears stale checkpoint entries for rules whose when gate recorded false (they are skips, not failures — issue #690); status reports them as skipped, not failed |
--profile |
— | — | Execution profile name, loaded from profiles/<NAME>.toml or profiles/<NAME>.oxoflow (see Execution profiles) |
--max-submitted |
— | profile default (50) |
Cluster jobs in flight at once (overrides the profile's max_submitted) |
--provenance |
— | — | Track output file checksums for later verification |
--arg |
— | — | Legacy form: set a workflow config value (KEY=VALUE). Repeatable. See [config] in workflow-format |
--bundle |
— | — | Execute from a published bundle (.tar.zst or .tar.gz). Extracts, verifies checksums, shows resource requirements, and prompts for confirmation |
--yes |
— | — | Skip the confirmation prompt when running a bundle (required in non-interactive sessions: CI, scripts, redirected input, or --json) |
--ai-recover |
— | — | Enable AI error recovery on rule failure |
--ai-max-retries |
— | — | Maximum AI attempts for recovery when a call fails (default: 1; only a failed call is retried) |
--samples |
— | all discovered samples | Sample selection: @path replaces the workflow's samples from a samplesheet, +@path appends (same-name groups merge, new groups added); names filter (or declare when the workflow ships no samples), first:N (pilot) and ready (samples whose entry inputs are complete) filter. Repeatable, comma-separated |
--rerun |
— | — | Force re-execution of this run's rules (ignore up-to-date checks). Checkpoint records for rules outside this run are kept |
--no-report-snapshot |
— | — | Skip the automatic report snapshot written after the run (see Report snapshots); resume has the same flag |
--background |
— | — | Detach the run into a background process and exit 0 immediately (see Background runs (--background)) |
--verbose |
-v |
— | Enable debug-level logging |
Examples#
Run with auto-discovery (when only one .oxoflow file exists)#
Run with main.oxoflow (priority discovery)#
Run with explicit workflow#
Run a workflow from a repository (nextflow-style)#
A public git repository can be executed directly — no clone, no bundle:
# Default branch — the gh: prefix is optional for github.com repositories
oxo-flow run owner/pipeline
# Pinned to a git tag or branch (@ref selects a git ref — unlike `pull`,
# where @tag means a GitHub Release bundle)
oxo-flow run owner/pipeline@v0.23.2
# The explicit forms work the same way (gh: can always be kept)
oxo-flow run gh:owner/pipeline
oxo-flow run gh:owner/pipeline@v0.23.2
# Any git URL or local repository directory
oxo-flow run https://example.com/team/pipeline.git
# A local checkout via file:// (must be an existing directory)
oxo-flow run file:///path/to/pipeline
A bare owner/pipeline argument is interpreted as a GitHub repository
unless it names something real on the filesystem: a path that exists, one
starting with ., or one ending in .oxoflow is always treated as a
local workflow path. Only a single-slash owner/name — optionally with
@ref and a trailing .git — maps to https://github.com/owner/name.git.
(So a missing data/pipeline.oxoflow keeps its normal "workflow file not
found" error instead of attempting a clone.) An https://…/*.git URL runs
that repository directly; file://<dir> runs an existing local checkout.
The repository is checked out into .oxo-flow/repos/<host>-<owner>-<name>
under the current directory (e.g. .oxo-flow/repos/github.com-owner-pipeline;
reused on later runs — delete the directory to force a fresh clone; the host
and owner prefixes keep same-named repos of different owners or hosts from
sharing a checkout), and the workflow file is auto-discovered
(main.oxoflow first). Because the clone is a read-only cache, the working directory
defaults to the current directory for repository runs — outputs, the
checkpoint, and the workdir lock all land next to your data, not inside the
clone. All other run semantics apply unchanged: --samples, --rerun,
config overrides, --workdir, checkpoint-aware dry-run previews, etc.
For github.com clones the official URL is tried first, then the
ghfast.top / gh-proxy.com mirrors automatically (see
China Mirrors).
Custom working directory#
Relative paths in the workflow resolve against the workflow file's
directory — the single resolution base (issue #427). When --workdir
differs from the workflow directory, the engine anchors workflow-shipped
files so they stay valid from the execution cwd:
- Environments: relative
conda/mamba/pixi/venv_requirementsfile specs resolve against the workflow directory. - Rule scripts: workflow-relative script files render absolute (workflow-rooted) so rule shells — which run with cwd = the workdir — can execute them.
- Workflow source files (zero duplication): declared rule inputs that
no rule produces, plus interpreter script paths (
python scripts/x.py), are materialized in the workdir as symlinks into the workflow directory at run start — the repo keeps the single copy of every source file. A pre-existing byte-identical workdir copy is replaced by a symlink (with a notice); a copy that differs is kept for the run and warned about (it may be a deliberate override). Repo paths a shell command mentions but no rule declares are reported as a note — declaring such a path as a rule input links it automatically. - Tool index sidecars (zero duplication): companion index files that
tools auto-derive from a linked file's path by convention —
samtools faidx/picard/GATK (ref.fa.fai,ref.fa.dict), bwa (ref.fa.amb/.ann/.bwt/.pac/.sa), minimap2 (.mmi), and the bgzip/tabix/htsfile index forms (.gzi,.tbi,.csi,.crai,.bai) — are linked alongside the file itself. Those paths never appear in a shell command, so static analysis alone cannot see them; the engine links every such sidecar that exists next to a linked workflow file, unless a rule produces the sidecar itself (a dedicated indexing rule keeps ownership). Tools driven through a bare index prefix (bowtie2-x index) should declare their index files as rule inputs directly. {input}and{log}: stay workdir-relative. Inputs are addressed from the run cwd (freshness/optional-input gates probe workdir-relative paths), and the engine creates log parent dirs under the workdir. Place workflow-consumed data under the workdir (or pass absolute paths).- input_groups disk scans: relative patterns scan the workflow
directory and emit workflow-absolute paths into
{input}. - Reference sources:
[references].sourceresolves against the workflow directory; reference outputs still land in the workdir. - Outputs: stay workdir-relative — they are run artifacts and land
under the workdir, matching the checkpoint and
.oxo-flowstate.
This keeps a workflow self-contained: for a shared, read-only workflow (a
fixed reference resource), point --workdir at the analysis directory
instead:
# Workflow lives in a central, read-only location; data and results
# belong to the current analysis directory
oxo-flow run /opt/pipelines/wgs.oxoflow --workdir .
# Outputs, .oxo-flow checkpoint, and lock all land in the workdir;
# dry-run/clean/report/resume accept --workdir too, and resume re-uses
# the workdir recorded in the checkpoint
Parallel execution#
Keep going on failure#
Retry failed jobs#
Run specific targets#
Target names may be prefixes: -t align matches every rule whose name
starts with align (including expanded instance names like
align_cohort_S1). The closure walks the instantiated DAG — it
includes each target plus every upstream rule that survives the plan-time
when gating, and skips when-gated variants entirely: a variant whose
when is false under the effective config never enters the execution set,
its upstream is not pulled in, and consumers of its (never-produced)
outputs are pruned too. Naming a when-gated-false target explicitly
reports it and aborts with "nothing to run" instead of planning a run
that cannot produce output.
Set a per-job timeout#
Limit resource usage#
# Use only 16 threads and 32GB memory
oxo-flow run pipeline.oxoflow --max-threads 16 --max-memory 32768
Cache environment setup#
# Cache environment setup state for faster subsequent runs
oxo-flow run pipeline.oxoflow --cache-dir .oxo-flow/cache
The cache would otherwise grow without bound, so after each run files
untouched for 90 days are removed (the next run starts clean). The age
limit is configurable per workflow — set cache_max_age_days = 0 to disable
aging entirely, or any other value to change the window:
Skip environment setup (when environments are pre-built)#
Backend preflight: fail fast on a missing environment backend#
After DAG construction and before any rule executes, run validates every
unique environment spec among the pending rules — once per spec (rule
templates fan out one spec across N samples, so the spec set is tiny even
for large cohorts). When a required backend binary is missing
(conda/mamba/pixi/docker/singularity/modules), the run aborts
before anything executes:
2 environment backend(s) required by pending rules are unavailable; no rules were run:
- conda: conda is not installed or not in PATH — required by 3 pending rule(s), e.g. `bwa_align` (environment: envs/alignment.yaml); install conda and ensure it is on PATH
- pixi: pixi is not installed or not in PATH — required by 1 pending rule(s), e.g. `report` (environment: pixi.toml); install pixi (https://pixi.sh) and ensure it is on PATH
Rules sharing one broken spec, and distinct specs failing with the
identical backend-missing message, collapse into a single entry — one root
cause per line, with the pending-rule count, an example rule, the primary
spec, and a mode-aware remediation hint. Nothing ran, so the --json
summary still reports the run as failed with zero counts (issue #142 H6).
See Troubleshooting — Environment Issues
for per-backend checks.
The remediation depends on why the spec failed (issue #848):
- Manifest file not found (pixi declared by relative path): the spec
is resolved against the workflow directory (the same base the executor
uses, issue #427), and the hint says to create the manifest at the
resolved path or correct the rule's
pixipath — installing the backend would not help when the file itself is absent. Other backends keep the install advice (a missing conda YAML surfaces only at task execution). - Backend binary missing (the manifest exists, or the backend cannot be
located for another reason): the hint is the per-backend install advice
(
install pixi (https://pixi.sh) and ensure it is on PATH, and so on).
This is the run-time safety net. validate and plan/dry-run preflight
pixi manifests up front (error E019) so the failure surfaces before you
commit to a run; run keeps the check because manifests can also be
created between validate and run.
Pass config values#
Every key in the [config] section of the workflow (including declarative
key = { default = …, required = …, … } entries) can be overridden from the
CLI — no extra flags to declare:
# Direct flag form
oxo-flow run pipeline.oxoflow --database refs/nt
# Attached form
oxo-flow run pipeline.oxoflow --database=refs/nt
# Key=value form directly after the workflow file
oxo-flow run pipeline.oxoflow database=refs/nt threshold=1e-3
# Legacy --arg form (still supported)
oxo-flow run pipeline.oxoflow --arg database=refs/nt --arg threshold=1e-3
# Empty string = explicit derive-if-empty override
oxo-flow run pipeline.oxoflow gene_bed=
oxo-flow run pipeline.oxoflow --gene_bed ""
Empty values select the derive-if-empty branch
An empty value is accepted and overrides the config key with the
empty string. Pipelines that auto-derive a value branch on
[ -n "{config.gene_bed}" ] — passing gene_bed= forces that branch
off (e.g. to re-derive from a new reference instead of reusing a
declared default path). Only the key must be non-empty; =value
alone is a syntax error.
Ordering: run flags before overrides
run flags (--json, --rerun, --samples, -j, …) must come
before positional KEY=VALUE overrides — the override list is
trailing, so a flag typed after it is reported with actionable guidance:
oxo-flow run pipeline.oxoflow min_quality=30 --json
'--json' is a command flag, not a config override.
Command flags must come before KEY=VALUE overrides, e.g.:
oxo-flow run <workflow.oxoflow> --json min_quality=30
For a config key that itself starts with dashes, use --arg KEY=VALUE
(also placed before positional overrides).
Execution profiles#
--profile <NAME> applies a reusable config supplement to a run. Profiles
are plain TOML (or .oxoflow) files placed in the workflow's own
profiles/ directory; .toml is tried before .oxoflow:
The name is just a filename — it carries no built-in meaning. What the
profile contains decides its effect: a [config]/[defaults] block
supplements config, a [cluster] block additionally routes execution to
a scheduler (see Cluster submission). Naming a
profile conda or slurm is a convention, not a mechanism — compute
environments are declared per rule via [rules.environment] and are
independent of profiles.
The profile's [config] table fills in keys the workflow does not set
itself by default — values already present in the workflow are never
overwritten. This lets one workflow carry per-environment fallback
defaults (laptop, cluster, cloud). Declare profile_mode = "override"
in the workflow's [workflow] section when a profile should instead
replace workflow values (deep-merged: nested tables merge recursively,
scalars and arrays are replaced) — the classic "cluster profile switches
threads/memory" case:
If the named profile file does not exist, run (and dry-run) fail
loudly with a nonzero exit and an error listing every available profile
in the profiles/ directory — a typo'd --profile must never silently
run with the workflow's own config. The same hard error applies when the
profile file exists but its TOML fails to parse or merge.
Cluster submission#
A profile that also carries a [cluster] block makes run --profile submit
to a scheduler instead of executing locally:
# profiles/slurm.toml
[cluster]
backend = "slurm" # slurm | pbs | sge | lsf — required
partition = "compute"
account = "lab01"
walltime = "24h" # a rule's own time_limit wins
max_submitted = 50 # jobs in flight at once
poll_interval = "30s"
extra_args = ["--exclusive", "--constraint=haswell"] # verbatim directives
[config]
reference = "/data/ref/hg38.fa"
Keys the profile omits fall back to the driver defaults: max_submitted
50, max_array_size 1001, poll_interval 5s. extra_args entries are
emitted as scheduler directives verbatim and are not validated — a typo
reaches the scheduler as written.
oxo-flow run pipeline.oxoflow --profile slurm
# one-off override of the profile's queue cap
oxo-flow run pipeline.oxoflow --profile slurm --max-submitted 100
Profiles are looked up in the workflow's own profiles/ directory —
there is no user- or system-level profile location, so a site profile shared
across a lab is copied into each workflow (or kept in the workflow
repository) rather than installed once per machine.
Both conditions are required: a [cluster] block must be in effect and
a profile must be named. A profile with only [config] keeps the local
path, and a workflow carrying its own [cluster] block still runs locally
until you pass --profile, so adding one never changes an existing run.
Rules become jobs after wildcard expansion, so a 100-sample scatter is 100
jobs, and dependencies are chained per instance. The run submits at most
max_submitted jobs at a time, tracks them to completion, and writes the
checkpoint exactly as a local run does — status, resume, and re-run
invalidation behave identically, and a second run --profile slurm with
nothing changed submits nothing.
Job arrays#
Same-template instances that are ready together and agree on every scheduler-visible directive (threads, memory, time limit, partition, account, GPU spec, environment, workdir) are grouped into one scheduler array automatically (issue #74):
[cluster]
backend = "slurm"
max_array_size = 1001 # optional — the scheduler's MaxArraySize; larger
# scatter groups are chunked into several arrays
- One submission instead of N
sbatch/qsubcalls — one queue entry that expands internally (SLURM--array=1-N, PBS-J, SGE-t, LSF-J "name[1-N]"). - Array elements are tracked individually: per-instance job directories
stay greppable (
jobs/<instance>/job.sh+job.id+status.json), andindex.jsonin the run directory maps each array index back to its instance, so the array never leaks into the human-facing layout. - Instances that differ on any directive fall out of the group and submit as single jobs — heterogeneity degrades gracefully.
- Arrays are transport-level: the JobRecord set, checkpoint, and resume semantics are identical to per-job submission.
Element-wise aftercorr chaining is not part of this slice: downstream
instances submit in ready batches as array elements finish, which is
correct but not as queue-efficient as true element-wise chaining.
Partial module runs (--module)#
A composed workflow (clinical pipelines like clindet: several
[[include]] modules chained together) can run just ONE module and the
producers of its declared inputs:
# run only the germline module (rules/20_germline.oxoflow) of a composed
# clinical pipeline — the module name defaults to the include's file stem
oxo-flow run main.oxoflow --module 20_germline
# multiple modules + rule targets union together
oxo-flow run main.oxoflow --module 20_germline --module 30_vcf_norm -t report
The closure is: the module's rules + every host rule producing one of its
declared concrete contract inputs; upstream DAG dependents come through
the regular target machinery. Unknown module names fail with the known
module list. dry-run --module previews the same set.
Run directory layout#
Each run leaves a directory you can navigate with ordinary tools:
.oxo-flow/runs/2026-08-17T14-30-05/ # `latest` symlinks to the newest
events.jsonl # append-only submit/complete/fail log
index.json # array index → instance name
jobs/<rule>/job.sh # the exact script submitted
jobs/<rule>/job.id # scheduler job id
jobs/<rule>/status.json
Every events.jsonl line carries an RFC 3339 ts, so the log reconstructs
a timeline rather than just an ordering:
{"ts":"2026-08-17T14:30:05.114Z","t":"SUBMITTED","rule":"align_S1","job":"4812345"}
{"ts":"2026-08-17T15:02:44.907Z","t":"COMPLETED","rule":"align_S1","job":"4812345"}
What a finished job records#
When a job leaves the queue, oxo-flow reads the scheduler's accounting store
(sacct, qstat -x -f, qacct, bacct) and writes what it found into
status.json:
{
"state": "COMPLETED",
"job_id": "4812345",
"command": "bwa mem ref.fa data/S1.fq > aln/S1.bam",
"submitted_at": "2026-08-17T14:30:05.114Z",
"finished_at": "2026-08-17T15:02:44.907Z",
"queue_wait_secs": 65,
"exit_code": 0,
"elapsed_secs": 1894,
"max_rss_mb": 24680,
"cpu_seconds": 14203
}
queue_wait_secs is submit-to-finish minus the scheduler's own elapsed
time — the driver never observes the moment a job starts, so the split
between waiting and working is only available by subtraction. The same
elapsed_secs is what benchmarks.wall_time_secs records in the
checkpoint, which therefore measures runtime rather than runtime plus queue
wait, and max_rss_mb / cpu_seconds populate the same benchmark fields a
local run fills from its own sampler.
How much of this appears depends on the site. Accounting is a deployment
choice: a SLURM cluster without slurmdbd answers sacct with nothing, and
LSF's bacct columns vary too much between versions to parse blind, so LSF
records state only. Fields the store did not report are omitted rather than
guessed — including exit_code, which stays absent for a failed job whose
real code could not be read, instead of defaulting to 1.
Outputs, the working directory, and logs are assumed to live on storage shared between the submitting host and the compute nodes.
Not yet implemented. The driver polls until every job reaches a terminal state; it does not yet re-attach to jobs still in flight if the driver exits, and Ctrl-C does not yet cancel submitted jobs. Until then, use
oxo-flow cluster status(no ids: lists your queued jobs) andoxo-flow cluster cancel <JOB_ID>...to inspect or clean up an interrupted run.
Cluster scheduler submission (SLURM, PBS, SGE, LSF) is configured
separately — see the cluster command.
Execute a published bundle#
# Run from a bundle (extracts, verifies checksums, shows resource requirements,
# then prompts for confirmation — requires an interactive terminal on both
# stdin and stderr)
oxo-flow run --bundle pipeline-bundle.tar.zst -j 16
# Skip confirmation for CI/scripts (also required when stdin is redirected
# or with --json, which is always non-interactive)
oxo-flow run --bundle pipeline-bundle.tar.zst -j 16 --yes
# .tar.gz format also supported
oxo-flow run --bundle pipeline-bundle.tar.gz -j 16 --yes
# Pull from remote and execute in one step
oxo-flow pull gh:user/repo@v1 && oxo-flow run --bundle repo-bundle.tar.zst -j 16 --yes
Pilot runs and scale-up#
For a large cohort, run a fast pilot on a subset first; the checkpoint
then skips the pilot samples when you scale up. --samples needs a
workflow that has a sample dimension (sample_pattern,
[[sample_groups]], sample_groups_file, or [[pairs]]): a
workflow with none rejects the flag with
--samples matched no samples in this workflow.
# Pilot: run the full pipeline on the first 2 samples only
oxo-flow run pipeline.oxoflow --samples first:2
# Or name the pilot samples explicitly (combines with first:N)
oxo-flow run pipeline.oxoflow --samples first:2,NA12899
# Scale up: no flag needed — completed samples are skipped automatically
oxo-flow run pipeline.oxoflow
# Preview the pilot plan without executing — with a checkpoint present,
# dry-run also predicts what will actually re-run (and what stays
# protected); see dry-run's "Checkpoint-Aware Rerun Preview"
oxo-flow dry-run pipeline.oxoflow --samples first:2
# Fix a config bug found by the pilot — only affected rules re-run
# automatically (see "Config changes" below); use --rerun to force everything
oxo-flow run pipeline.oxoflow --rerun
--samples filters every sample source (sample groups, sample_pattern
discovery, and experiment/control pairs). --rerun re-executes the rules
selected for this run even when outputs are up to date — combine it with
--samples to re-run just the pilot subset while keeping the other
samples' completed records.
After a --samples run, a pilot summary is printed: samples run,
wall time, per-sample time, and a linear projection for the full cohort.
Scientific-preflight findings are reported as distinct findings with
their instance counts — one template rule appearing once per sample shows
as 1 distinct finding (3 instance(s)), not 3 finding(s). When the
workflow enables [ai], a plain-language pilot report (health assessment
and scale-up advice) is appended automatically.
Incremental data arrival: --samples ready#
Sequencing centers deliver data in batches, so a cohort's fastq files trickle in over days. Instead of waiting for every sample, run the analysis as data arrives:
# Which samples can be processed right now, and what is still missing?
oxo-flow dry-run pipeline.oxoflow
# Sample readiness: 87/100 complete, 13 waiting
# ⏳ NA12891 (missing: data/NA12891_R2.fastq.gz), …
# Run only the samples whose entry inputs are complete
oxo-flow run pipeline.oxoflow --samples ready
# … more data arrives …
oxo-flow run pipeline.oxoflow --samples ready
# Done: 5 succeeded, 87 skipped — the checkpoint skips completed samples
A sample is ready when every external input belonging to it exists —
that is, every rule input (after wildcard and {config.x} expansion) that
the workflow itself does not produce. Intermediate products are never
checked (producing them is the DAG's job), and optional = true rules do
not block readiness (the executor skips them when their inputs are absent).
Relative paths resolve against the workflow file's directory (or
--workdir) — the same place rules actually run from.
Semantics:
readyis a special--samplesvalue, not a new syntax: it resolves to the names of ready samples and combines withfirst:Nand explicit names as a union. Because it is reserved, a sample literally namedreadycannot be selected by name.- With zero ready samples
runaborts and lists the waiting samples;dry-run --samples readyinstead reports the full cohort state. - For
[[pairs]]workflows, a pair is kept only when both experiment and control inputs are complete; otherwise it is skipped with a note. Missing files that belong to no specific sample (shared references) are reported as workflow-level inputs. - The report is also available as JSON via
dry-run --json(thesamplesblock), and--samples readyworks intest --runas well.
Background runs (--background)#
run --background (and resume --background) detaches execution into a
background process and returns to the shell immediately:
# Foreground prints one line and exits 0; the run continues in the background
oxo-flow run pipeline.oxoflow --background
# started in background (pid 48291) · log: .oxo-flow/logs/oxo-flow.log ·
# monitor: oxo-flow status .oxo-flow/checkpoint.json · stop: kill 48291
# Detach a resumed run the same way
oxo-flow resume .oxo-flow/checkpoint.json --background
Mechanics:
- The foreground process spawns a detached child that re-runs the same
command with
--backgroundremoved — every other flag (-j,--profile,--samples,--rerun,--log-file, …) passes through verbatim, so a background run behaves exactly like a foreground one: checkpoint, workdir lock, report snapshots, and resume semantics are unchanged. - The child's stdout and stderr are redirected to the run log
(
--log-fileif given, otherwise.oxo-flow/logs/oxo-flow.login the workdir). In this mode the tracing tee is skipped on purpose — the redirect already captures everything, and a second writer would duplicate every line (issue #194 A3). - The run log archives BOTH the structured tracing stream AND the
user-facing progress narrative (
Running:/✓/Done:) plus one JSON line per execution event (workflow_started,rule_started,rule_completed,workflow_completed) — the whole run is replayable from the single file (issue #194 B1/B3). - The child's pid is written to
.oxo-flow/background.pidin the workdir (also shown in the summary) while the run is in flight; the file is removed when the child exits, so a stale pid never survives a finished run. - The child gets its own process group, so it survives terminal close and Ctrl-C at the shell — it keeps running until it finishes or is killed.
- The foreground exits 0 once the child is spawned; it does not wait. If the spawn fails (for example because another run already holds the workdir lock), the foreground exits nonzero with the reason.
Which commands accept --background:
runandresume— the only long-lived terminal-bound commands. Cluster runs go throughrun --profile <name>and therefore detach the same way:run --profile slurm --backgroundsubmits and polls the cluster jobs entirely inside the background process.dry-runintentionally does not: a preview is fast and read-only — keep it in the foreground.- Other commands (
pull,test,export,ai, …) are either quick or one-shot; run them before the background run and keep them foreground.
Monitoring a background run works exactly like a foreground one: poll
oxo-flow status .oxo-flow/checkpoint.json (or status --timing), read
the run log, and check the report snapshots in .oxo-flow/reports/. Stop
a background run with kill <pid> (Unix) or taskkill /PID <pid> —
the pid is in the summary line and (while the run is in flight) in
.oxo-flow/background.pid.
Notes:
- The global
--jsonis rejected with--background: the foreground only launches the detached child, so there is no run summary to emit. Drop--jsonor run in the foreground to get the machine-readable summary. - Bundle runs need an explicit
--workdir(the bundle's extracted directory is created per-process, so it cannot be predicted from the foreground invocation) and--yes— a detached child cannot prompt.--workdiris also what bundle discovery consults as a fallback: asample_pattern/pairs_patternthat matches nothing inside the extraction is retried against the workdir (#751). --backgroundis a launcher flag, not a workflow setting: the child re-parses the remaining argv, so flag ordering rules (command flags beforeKEY=VALUEoverrides) apply to the whole command line as usual.
Checkpointing and Resuming#
oxo-flow automatically persists execution state to a checkpoint file after every rule completion. This enables:
- Resuming failed runs: If a job fails or is interrupted, simply run
oxo-flow runagain. The engine will skip all rules already marked ascompletedand resume from the first pending task. - State Inspection: Use the
oxo-flow statuscommand to view the progress of a run from its checkpoint file.
Metadata Directory#
By default, checkpoints are saved in a hidden .oxo-flow/ directory located in the same folder as the workflow file.
- Filename:
checkpoint.json(the name is always the same regardless of workflow name)
Rule run records in the checkpoint#
Each executed rule carries a rule_runs entry in checkpoint.json with the
expanded command, the process exit code, a bounded stdout/stderr tail, the
resolved
report caption, and (since the cancellation-semantics change) the terminal
status ("failed", "cancelled", "timed_out", …), the
skip_reason when the rule did not run to completion, and the
terminating signal when the process died to one. The three newer
fields are optional and absent in checkpoints written by older engines —
consumers must infer legacy records as before (failure when the rule is in
failed_rules).
Timestamp semantics (issue #850): the engine stamps
enqueued_at when an execution attempt is submitted and stamps
started_at only after the rule has acquired its resources — the
moment actual execution begins. Their difference is the time the rule
spends waiting in the resource pool (same split as the cluster driver's
queue_wait_secs), and all wall-time reporting (duration_ms, the
report's Benchmarks Wall Time column, benchmarks.wall_time_secs) is
measured from started_at, so a rule parked behind busy resources never
inflates its reported runtime. These two timestamps live on the engine's
in-memory job records and are not part of rule_runs entries; the
checkpoint persists the resulting wall time via benchmarks.wall_time_secs,
and cluster runs additionally persist the split per job as
queue_wait_secs in jobs/<rule>/status.json (see
What a finished job records).
The tails are truncated to the last 64 KiB per stream by default and
prefixed with a … marker when cut. Set OXO_FLOW_OUTPUT_TAIL_BYTES to
change the size (bytes; an invalid or zero value falls back to the default
with a warning), e.g. OXO_FLOW_OUTPUT_TAIL_BYTES=10485760 (10 MiB) for
error-heavy tools whose message scrolls past the tail. The in-RAM capture
window widens to match when the setting exceeds its 1 MiB floor — budget
roughly 2 × (1 MiB + tail) of RAM per running job. Tails are
bytes-bounded, so a multi-byte UTF-8 cut lands on the nearest character
boundary (rounding down); older checkpoints that only carry the legacy
2048-character tails remain readable (the new size applies from the next
recorded run onward).
Declared log files vs captured tails#
A rule's declared log path is separate from this capture: the engine
creates the log's parent directory, but the file only gains content when
the rule's shell redirects into it (2> {log}) and the tool writes to
stderr. Many bioinformatics tools self-log to their own location instead
(e.g. a logs/<date>.txt file they manage themselves), leaving the
declared log empty — an empty declared log never means no diagnostics
were captured; the tails above are persisted in rule_runs regardless.
When a rule succeeds, its declared log is empty (0-byte file or empty
directory), and both captured tails are empty too, the run emits one
Info-level hint that the tool likely self-logs elsewhere — quiet tools
with output in any of the three stay silent (issue #832).
Workflow version in the checkpoint#
When the workflow file lives inside a git repository, oxo-flow records the
repository's HEAD commit SHA in the checkpoint (workflow_git_sha) at
run start. Every result set is therefore auditable to the exact workflow
version that produced it: oxo-flow provenance verify prints the SHA, and
report snapshots carry it through the checkpoint they embed.
The lookup is best-effort — running outside a git repository simply omits the field and never fails the run. Commit your workflow before running to get fully version-audited results. See Workflow Versioning for the full model.
Run logs#
Every run — and resume, which re-enters the same path — archives its own
log under the workdir:
.oxo-flow/logs/oxo-flow.logis always the latest run; on each new run the previous log rotates to.oxo-flow/logs/oxo-flow.log.1, then.2, … up to.9(the oldest backup is deleted).--log-file PATHoverrides the location; the same rotation applies toPATH.- The log header names the exact workflow version that produced the record:
timestamp, oxo-flow version, full command line, workflow path,
workflow_name,workflow_version,git_sha(the repository HEAD recorded by the run — see Workflow Versioning), and workdir. - The engine's tracing stream (rule start/finish verdicts, warnings, errors) is written into the log alongside stderr. Logging is best-effort: if the file cannot be opened or written, the run continues with stderr-only logging and a warning. Dry-run remains stderr-only.
Report snapshots#
After every run (and resume) — unless --no-report-snapshot is given —
oxo-flow writes a JSON report snapshot of the final checkpoint:
.oxo-flow/reports/report-<UTC timestamp>.json— the full report data model, so a run leaves a machine-readable report behind with no reporting step needed (a-Nsuffix is used when two snapshots land in the same second).oxo-flow/reports/index.json— a JSON array of{generated_at, workflow, checkpoint, report}entries, kept sorted bygenerated_at; the last entry is the newest snapshot
Snapshot failures are warnings — a reporting hiccup never fails a run.
oxo-flow report with no -o writes auto-discovered reports into the
same .oxo-flow/reports/ directory.
Config changes and precise invalidation#
The checkpoint records a snapshot of the effective config values and a structural fingerprint of every completed rule. On each run, oxo-flow compares them against the current workflow:
- Changed config keys invalidate exactly the rules that reference them plus their DAG downstream; unrelated completed rules keep their checkpoint records. How a rule references a key decides the verdict:
- A key interpolated into
shell,script,input/outputpaths,envvars, orparams({config.<key>}) always invalidates — the new value bakes straight into the command or paths. - A key referenced only inside a
whencondition invalidates only when the condition's truth value actually flips under the new config. A toggle that leaves the gate true (or false) — e.g.(config.refine_bins || config.run_checkm)switching which term is true — reuses the completed rules instead of re-running hours-long chains whose inputs and commands are identical (issue #198). Flipping a gate in either direction still invalidates: the engine pre-marks completed rules without re-evaluatingwhen, so only plan-time detection can retire a now-skipped producer or cascade a newly-activated one to its consumers. - Edited rule definitions (shell, inputs, outputs, environment, …) are
caught by the per-rule fingerprint and invalidate the same way. A
file-backed environment spec (pixi manifest, conda/mamba YAML, venv
requirements) is fingerprinted by content, not just by the path the
rule names — pixi specs also cover the sibling
pixi.lock, since a lock re-pin changes exactly which packages the environment materializes. Editing any of these files in place invalidates the rules that reference them. Checkpoints written before the content tagging lack these digests, so the first run after that upgrade re-runs rules with file-backed environments once (the same safe-but-costly bootstrap as the config-tracking baseline below). samples_list/samples_<group>are engine-injected and never trigger invalidation, so toggling--samplesbetween runs stays cheap. This covers rules whose input list is baked from an injected key (expand_inputsoverconfig.samples_list, a gather-rule pattern): their rule fingerprint and input manifest change with the selection, and oxo-flow treats that change as non-invalidating. When such a rule is skipped because of a selection change,runprints a warning naming the rule — its outputs still reflect the previous run's full sample set, and--rerunis the documented way to regenerate them under the new selection. The exclusion is narrow and verified: the rule's other definition fields must be byte-identical, and the current input file set must be exactly reproducible from the injected value with unchanged content — any genuine edit still invalidates. Checkpoints written by older binaries lack the input-excluded fingerprint, so the first subset run after an upgrade may re-run such a rule once (safe default).
# min_quality is interpolated by fastp_trim; only it and its downstream re-run.
# A flag toggled in a when-only rule whose gate stays true re-runs nothing —
# the run reports those rules as reused instead.
oxo-flow run pipeline.oxoflow min_quality=30
Config change:
min_quality: 20 → 30
when-condition unchanged, reused: bin_classify
→ invalidated 3 (1 directly affected), re-running 3/10 this run, skipping 7
⊝ index_reference (already completed)
Running: fastp_trim
…
When a when gate closes at execution time, the rule is reported under
Skipped (when condition false) with a ⊘ marker (or ⊝ in the
config-change reuse view) — a gated-off rule is a skip, never a failure
(see status for the checkpoint-side
display, including legacy checkpoints that carry stale failed_rules
entries).
Rules changed by a transformed split set (transform.split.values_from)
are detected through their baked input lists, so chunked rules re-combine
correctly when the split values change.
Limitations (correctness-first, documented deliberately):
- Changing
threads/memory(performance knobs) does not invalidate results; changing anything else in a rule — including shell comments — does. - A checkpoint written before config tracking was introduced is adopted as
a one-time baseline: the first post-upgrade run reuses everything and
records the snapshot; changes made before that baseline cannot be
detected. The same applies to the per-rule
whenverdicts behind gate-aware reuse: a pre-verdict checkpoint invalidates when-referencing rules once on the first changed-key run, records their verdicts, and reuses stably from then on. - Config values declared
sensitiveare stored in the snapshot as SHA-256 digests, never as plaintext. - Concurrent
oxo-flow runinvocations on the same workdir are prevented by an exclusive lock on.oxo-flow/lock(held for the whole run). The second invocation fails with a clear error instead of silently racing oncheckpoint.json; the lock releases automatically when the first process exits (even on a crash).cleanrefuses to delete while a run is active unless--forceis given.
Input changes and manifest invalidation#
The checkpoint also records an input manifest for every completed rule: the file set its inputs resolved to at completion time, with each file's path, size, and modification time — plus a content hash for files up to 64 MiB. On every run, oxo-flow re-resolves the inputs and compares:
- Literal glob inputs (
data/*.txt) detect added and removed files — a new file matching the glob rebuilds the rule (and its DAG downstream). - Directory inputs (e.g.
input = ["results"]— a path that resolves to a directory) are listed recursively, so files added or changed anywhere inside invalidate the rule.Dirinputs with apatternfilter track only matching files. - Plain file inputs are content-addressed up to 64 MiB: a same-size
rewrite is detected even when the mtime is preserved, and a mere
touchno longer invalidates. Larger files compare size and mtime, closing the same stale-reuse hole for ordinary files. -
expand_inputs-baked inputs whose producer iswhen-gated off under the active config are tolerated as absent: the manifest records the existing subset instead of erroring, so a no-op rerun does not invalidate the rule and cascade its producers (issue #757). If a later config change produces the missing files, the manifest then mismatches and the rule invalidates — the flip is still caught. -
Missing external source inputs fail fast before any work starts: a required input that no rule in the DAG produces and whose concrete files are absent aborts the run before scheduling with
workflow is missing required source input(s) — no rule produces them, naming each affected rule and pattern (exit non-zero). Previously a fully-completed checkpoint plus a deleted raw data file could exit 0 from stale checkpoint credit ("re-running 0 producer rule(s)"); that silent-success hole is closed. The gate never fires for generated intermediates (they have producers and are governed by the manifest cascade below),[[references]]outputs (built later in the same run), inputs discoverable through the workflow'ssample_patterndirectory, when-gated-off rules, zero-fan-out scatter rules, oroptionalrules.
Invalidated rules bypass the executor's freshness gate and re-execute even when their outputs exist and look recent; downstream rules that consumed their outputs rebuild in the same run.
Limitations (correctness-first, documented deliberately):
- Detection is a hybrid policy: files up to 64 MiB are content-hashed
(
sha256), so a same-size rewrite is detected even when the mtime is preserved — and a meretouchno longer invalidates. Larger files (multi-gigabyte intermediates) keep the size+mtime policy, likemakeand Snakemake's default mode: hashing every input on every run would cost O(total input bytes), prohibitive for bioinformatics-scale files. A large file rewritten with identical size and a deliberately preserved mtime is not detected; use--rerunto force. Legacy checkpoints (entries written before hashing) keep comparing size+mtime instead of invalidating everything once. - Only one glob level is tracked per input pattern; brace-expansion
patterns (
*.{txt,log}) are not expanded by the manifest scanner, so their matched set is not tracked (the shell still expands them at runtime). - Symlinked directories are recorded as single entries, never traversed (cycle-safe), so changes inside a symlinked directory are tracked at the symlink's own metadata level.
- Rules that clean up their inputs at the end of a successful run
(
transform.cleanup = truechunk consumers) are not manifest-tracked: their inputs are engine-managed intermediates governed by the upstream-cascade invalidation instead. - Optional rules with permanently-absent inputs (
optional = trueoroptional = "any", issue #633): when a declared input cannot be resolved because its producer was never instantiated (aninput_groupspattern that matched no files, or an endedness-filtered branch that never ran), the snapshot skips that entry and records a manifest for the inputs that do exist — the rule stays skipped/completed across runs instead of cascading a full re-run every invocation. The same grace applies during invalidation detection: a missing input with no producer in the DAG is absent by design for an optional rule, while a required rule with a missing input still invalidates (a deleted raw data file must re-run). A failed manifest snapshot is reported on stderr and in the run log (⚠ could not snapshot inputs of rule …) instead of being silently dropped. - A checkpoint written before input tracking was introduced adopts the current file set as a one-time baseline (the first post-upgrade run reuses everything); changes made before that baseline cannot be detected.
Temporary rules (temporary = true)#
Pipeline intermediates are usually the largest files in a run (per-sample
BAMs, unsorted alignments) and often useless once every downstream rule has
consumed them. Marking a rule temporary = true deletes its outputs after
a fully successful run — once every dependent has completed — and
records a tombstone in the checkpoint:
This is checkpoint-aware, not a blind delete:
- A plain re-run skips the rule like any completed rule — the output is not regenerated just because it is missing.
- When a dependent actually needs the outputs again (its own outputs were deleted, its inputs changed, or the config changed), the producer is regenerated first (lazy cascade-up), then the dependent runs — the same order as the original run.
- Failed runs keep the outputs (the deletion only happens when nothing failed), so debugging a broken run never has to re-fetch intermediates.
- Leaf rules (no dependents) keep their outputs — there is no downstream work for them to enable.
↻ temporary outputs needed again — re-running 1 producer rule(s): trim_S1
…
⊘ temporary outputs deleted for 'trim_S1' (1 file(s), regenerated on demand)
Dry-run predicts exactly when a tombstoned rule will regenerate
([rerun: upstream of X]) — see
Checkpoint-aware rerun preview.
Disk-pressure monitoring ([engine])#
Large pipelines can fill a disk mid-run. Declaring an [engine] section in
the workflow makes the engine measure the working directory's free space at
startup, at every scheduling round, and every 60 s while work is parked:
[engine]
min_free_disk = "50G" # floor — entering Hold below it
reclaim_free_disk = "100G" # optional reclaim trigger; defaults to min_free_disk
The response ladder:
- Reclaim (
min ≤ free < reclaim): completed rules' temporary outputs are tombstoned and unlinked mid-run — the same checkpoint-aware reclaim astemporary = truecleanup, but while the run is still going. Only when every dependent is completed and its own outputs still exist (a dependent whose outputs are missing will re-run and needs its inputs). The narration line is⊘ reclaimed temporary outputs of '<rule>' …. - Hold (
free < min): new rule dispatches are withheld; in-flight rules continue. If nothing is in flight and nothing is reclaimable, the run parks and re-measures every 60 s. - Abort (Hold persists ~2 parked rounds ≈ 2 min): the run checkpoints
and fails. Free space manually (delete intermediates or
oxo-flow clean), thenoxo-flow resume— or lowermin_free_disk.
Protected outputs (protected_output) are never reclaimed, and the final
(leaf) rule's outputs always survive. See
[engine] in the workflow format reference
for the full field table and edge cases.
Tuning knobs (both fall back to their defaults when unset, invalid, or
0, so a typo can never disable the abort):
OXO_FLOW_DISK_TICK_SECS— seconds between measurements while parked (default 60).OXO_FLOW_DISK_HOLD_ROUNDS— grace park rounds under Hold before the run aborts (default 2).
Known boundaries:
- Mid-run reclaim of fan-out producers (
output_patternrules with per-instance values) does not fire — the event loop tracks reclaim candidates by rule name and holds no instance bindings, so those intermediates are only cleaned by the post-run sweep (instance-level, issue #844) or the next run. Fail-safe: no reclaim rather than a wrong one. - Free space is measured at the working directory's filesystem, which is volume-level (statvfs): thresholds apply to the whole volume, not to a quota.
- Keep scratch-heavy tools off the measured volume only if they are not cleaned by their rules; outputs declared by rules are always accounted for by the ladder.
Forcing Execution#
To bypass checkpoints and re-execute rules that have already completed, use
oxo-flow clean to remove outputs and the checkpoint file
before running, or --rerun to re-execute this run's rules while keeping
checkpoint records for rules outside the run.
Checkpoint re-entry (checkpoint = true)#
A checkpoint rule discovers new values at runtime: after it completes, the
engine reads its checkpoint_manifest, merges the declared samples, and
executes the newly created rule instances in the same run. The checkpoint
records each re-entry (reentries), so a resume replays it and reconstructs
the same plan; invalidating the checkpoint rule revokes its samples until it
re-runs. A missing or unparsable manifest fails the checkpoint rule. See
Workflow Format for
the config surface.
Output#
oxo-flow v0.23.2 — Rust-native bioinformatics pipeline engine
DAG: 5 rules in execution order
1. fastqc
2. trim_reads
3. bwa_align
4. sort_bam
5. call_variants
⠋ [00:15] [████████████░░░░░░░░] 3/5 ETA:0:00:42 (executing sort_bam)
✓ fastqc (2.1s)
✓ trim_reads (15.0s)
✓ bwa_align (42.3s)
✓ sort_bam (10.2s)
✓ call_variants (8.4s)
Done: 5 succeeded, 0 skipped, 0 failed
✓ 5 output files verified (118.3MB total)
A progress bar shows execution progress with:
- Elapsed time
- Current position / total rules
- Estimated time remaining (ETA), plus the rule currently executing
- The final summary line reports succeeded / skipped / failed counts
Notes#
- The workflow file is optional; if not specified, auto-discovery searches for
main.oxoflowfirst, then any*.oxoflowfile alphabetically - If no
.oxoflowfile is found, an error message suggests runningoxo-flow initto create one - The DAG is built and validated before any rules execute
- Rules are executed in topological order; independent rules may run in parallel up to the
-jlimit - If
--keep-goingis not set, execution stops at the first failure --keep-goingchanges scheduling, never the verdict: it runs the remaining rules past a failure, but the run still exits non-zero when arequiredrule failed — scripts and the web UI (which classifies by exit code) see the truth. Onlyrequired = falsefailures leave the exit code at 0.- The
--retryflag re-runs failed jobs up to N times before marking them as failed - A timeout of
0means no timeout - Resource constraints (
threads,memory) in rules are checked against available resources before execution - Pre-flight budget check: When
--max-threadsor--max-memoryis explicitly set, the engine checks all rules before execution starts. Any rule whose requirements exceed the budget is reported immediately (fast-fail), preventing mid-pipeline failures. An explicit budget is a hard limit. - Backend preflight: before any rule executes, every unique environment spec of the pending rules is validated once — a missing backend (
conda,mamba,pixi,docker,singularity,modules) aborts the run with one collapsed entry per root cause and "no rules were run" (see Backend preflight). - Auto-detected capacity is soft: Rules whose declared requests exceed the machine's detected threads/memory are not rejected — the declared request is the tool's upper bound (often an upstream HPC label), not a scheduling requirement. The engine warns and clamps the pool reservation, so an over-capacity rule runs alone (serialized) instead of blocking the workflow.
- Deadlock detection: If pending rules remain while nothing is running and none can become ready (typically an upstream failure), the engine reports
Deadlock detected: N rules stuckwith the stuck rule names. Resource waits cannot deadlock: over-capacity requests are clamped, and explicit budget violations fail fast before any rule runs. - Target-aware execution: The
-tflag supports prefix matching —-t almatches all rules whose names start with "al". Use this to run a subset of the workflow, similar tomake <target>. Only the named targets and their transitive upstream dependencies are executed; downstream rules are excluded. - Default targets: rules marked
target = trueare whatrun/dry-runbuild when neither-tnor--moduleselects anything — the marked rules and their transitive producers. With no marked rule the full DAG runs (backward compatible); an explicit-t/--modulealways wins. - Abort cancellation semantics: with fail-fast aborts, rules that were merely killed by the abort (signal death, no exit code) are recorded as cancelled, not failed — they are excluded from
failed_rulesand the failure counts, and narrated like skips. A sibling that failed on its own merit (exit code present) keeps its failure record. The checkpoint's per-rulerule_runsentries distinguish the two via theirstatus/skip_reason/signalfields. - Setting
--max-threads 0or--max-memory 0auto-detects system resources - Environment setup is performed automatically before first use of each environment (conda, pixi, docker, singularity, venv)
- Use
--skip-env-setupwhen environments are pre-built to avoid redundant setup - Use
--cache-dirto persist environment setup state across runs for faster startup