Skip to content

oxo-flow run#

Execute a workflow.


Usage#

oxo-flow run [OPTIONS] [WORKFLOW] [KEY=VALUE]...

Arguments#

Argument Description
[WORKFLOW] Path to the .oxoflow workflow file. Optional — if not specified, auto-discovery searches for: (1) main.oxoflow in current directory, (2) alphabetically first *.oxoflow file in current directory.
[KEY=VALUE]... Direct config overrides: KEY=VALUE and --KEY=VALUE work for any key; the --KEY VALUE space form also works for any key present in [config] (plain key = "value" entries included) — a key absent from [config] is a hard error (an unknown flag would otherwise be indistinguishable from a mistyped option; a typo'd --key=val after the workflow still sets a config var, so prefer the = forms for undeclared keys). An override key that differs from a declared key only by hyphens vs underscores is canonicalized onto the declared spelling — --api-token=x sets the declared key api_token (and inherits its sensitive masking); keys matching nothing declared keep the literal spelling

Options#

Option Short Default Description
--jobs -j 1 Maximum number of concurrent jobs
--keep-going -k — Continue execution when a job fails
--json — — Emit a machine-readable JSON summary on stdout after the run (command, status, results counts, resources). The document is emitted on every exit path — completed, failed, and aborted (preflight failures, budget breaches, cluster runs, plain-failure aborts) — so --json consumers never get zero bytes. status is "completed" only when the run finished with zero required failures; every other exit is "failed". Human progress and result output on stderr is not suppressed — this flag appends the summary, it does not switch the run to machine-only mode. resources lists every benchmarked rule (rule, status, wall_time_secs, peak_rss_mb, cpu_seconds, retries) — the same data as the report's Benchmarks table; it is empty on aborted paths and on cluster runs (use the report's sacct-based Resource Accounting section there)
--workdir -d Workflow file's directory Working directory for execution
--log-file — .oxo-flow/logs/oxo-flow.log in the workdir Write the run log to a custom path instead (relative paths resolve against the workdir; previous logs rotate to PATH.1 … PATH.9)
--target -t All rules Run only specific target rules
--module — — Run one include module and the producers of its declared inputs (repeatable; unions with --target). Module names are the include's name field or its file stem (see Partial module runs)
--retry -r 0 Number of times to retry failed jobs
--timeout — 0 (disabled) Timeout per job in seconds, or a duration like 1h/30m
--max-threads — 0 (auto-detect) Maximum CPU threads available for execution
--max-memory — 0 (auto-detect) Maximum memory in MB available for execution
--skip-env-setup — — Skip environment setup (assume environments are ready)
--skip-ref-build — — Skip automatic reference/index building (assume pre-built)
--cache-dir — — Directory for caching environment setup state (entries untouched for 90 days are cleaned up after each run; override with the cache_max_age_days config key, 0 disables aging)
--resume-failed — — Resume only failed rules from a previous run. Also clears stale checkpoint entries for rules whose when gate recorded false (they are skips, not failures — issue #690); status reports them as skipped, not failed
--profile — — Execution profile name, loaded from profiles/<NAME>.toml or profiles/<NAME>.oxoflow (see Execution profiles)
--max-submitted — profile default (50) Cluster jobs in flight at once (overrides the profile's max_submitted)
--provenance — — Track output file checksums for later verification
--arg — — Legacy form: set a workflow config value (KEY=VALUE). Repeatable. See [config] in workflow-format
--bundle — — Execute from a published bundle (.tar.zst or .tar.gz). Extracts, verifies checksums, shows resource requirements, and prompts for confirmation
--yes — — Skip the confirmation prompt when running a bundle (required in non-interactive sessions: CI, scripts, redirected input, or --json)
--ai-recover — — Enable AI error recovery on rule failure
--ai-max-retries — — Maximum AI attempts for recovery when a call fails (default: 1; only a failed call is retried)
--samples — all discovered samples Sample selection: @path replaces the workflow's samples from a samplesheet, +@path appends (same-name groups merge, new groups added); names filter (or declare when the workflow ships no samples), first:N (pilot) and ready (samples whose entry inputs are complete) filter. Repeatable, comma-separated
--rerun — — Force re-execution of this run's rules (ignore up-to-date checks). Checkpoint records for rules outside this run are kept
--no-report-snapshot — — Skip the automatic report snapshot written after the run (see Report snapshots); resume has the same flag
--background — — Detach the run into a background process and exit 0 immediately (see Background runs (--background))
--verbose -v — Enable debug-level logging

Examples#

Run with auto-discovery (when only one .oxoflow file exists)#

# No need to specify the workflow file
oxo-flow run

Run with main.oxoflow (priority discovery)#

# If main.oxoflow exists, it's automatically used
oxo-flow run

Run with explicit workflow#

oxo-flow run pipeline.oxoflow

Run a workflow from a repository (nextflow-style)#

A public git repository can be executed directly — no clone, no bundle:

# Default branch — the gh: prefix is optional for github.com repositories
oxo-flow run owner/pipeline

# Pinned to a git tag or branch (@ref selects a git ref — unlike `pull`,
# where @tag means a GitHub Release bundle)
oxo-flow run owner/pipeline@v0.23.2

# The explicit forms work the same way (gh: can always be kept)
oxo-flow run gh:owner/pipeline
oxo-flow run gh:owner/pipeline@v0.23.2

# Any git URL or local repository directory
oxo-flow run https://example.com/team/pipeline.git

# A local checkout via file:// (must be an existing directory)
oxo-flow run file:///path/to/pipeline

A bare owner/pipeline argument is interpreted as a GitHub repository unless it names something real on the filesystem: a path that exists, one starting with ., or one ending in .oxoflow is always treated as a local workflow path. Only a single-slash owner/name — optionally with @ref and a trailing .git — maps to https://github.com/owner/name.git. (So a missing data/pipeline.oxoflow keeps its normal "workflow file not found" error instead of attempting a clone.) An https://…/*.git URL runs that repository directly; file://<dir> runs an existing local checkout.

The repository is checked out into .oxo-flow/repos/<host>-<owner>-<name> under the current directory (e.g. .oxo-flow/repos/github.com-owner-pipeline; reused on later runs — delete the directory to force a fresh clone; the host and owner prefixes keep same-named repos of different owners or hosts from sharing a checkout), and the workflow file is auto-discovered (main.oxoflow first). Because the clone is a read-only cache, the working directory defaults to the current directory for repository runs — outputs, the checkpoint, and the workdir lock all land next to your data, not inside the clone. All other run semantics apply unchanged: --samples, --rerun, config overrides, --workdir, checkpoint-aware dry-run previews, etc.

For github.com clones the official URL is tried first, then the ghfast.top / gh-proxy.com mirrors automatically (see China Mirrors).

Custom working directory#

Relative paths in the workflow resolve against the workflow file's directory — the single resolution base (issue #427). When --workdir differs from the workflow directory, the engine anchors workflow-shipped files so they stay valid from the execution cwd:

  • Environments: relative conda / mamba / pixi / venv_requirements file specs resolve against the workflow directory.
  • Rule scripts: workflow-relative script files render absolute (workflow-rooted) so rule shells — which run with cwd = the workdir — can execute them.
  • Workflow source files (zero duplication): declared rule inputs that no rule produces, plus interpreter script paths (python scripts/x.py), are materialized in the workdir as symlinks into the workflow directory at run start — the repo keeps the single copy of every source file. A pre-existing byte-identical workdir copy is replaced by a symlink (with a notice); a copy that differs is kept for the run and warned about (it may be a deliberate override). Repo paths a shell command mentions but no rule declares are reported as a note — declaring such a path as a rule input links it automatically.
  • Tool index sidecars (zero duplication): companion index files that tools auto-derive from a linked file's path by convention — samtools faidx/picard/GATK (ref.fa.fai, ref.fa.dict), bwa (ref.fa.amb/.ann/.bwt/.pac/.sa), minimap2 (.mmi), and the bgzip/tabix/htsfile index forms (.gzi, .tbi, .csi, .crai, .bai) — are linked alongside the file itself. Those paths never appear in a shell command, so static analysis alone cannot see them; the engine links every such sidecar that exists next to a linked workflow file, unless a rule produces the sidecar itself (a dedicated indexing rule keeps ownership). Tools driven through a bare index prefix (bowtie2 -x index) should declare their index files as rule inputs directly.
  • {input} and {log}: stay workdir-relative. Inputs are addressed from the run cwd (freshness/optional-input gates probe workdir-relative paths), and the engine creates log parent dirs under the workdir. Place workflow-consumed data under the workdir (or pass absolute paths).
  • input_groups disk scans: relative patterns scan the workflow directory and emit workflow-absolute paths into {input}.
  • Reference sources: [references].source resolves against the workflow directory; reference outputs still land in the workdir.
  • Outputs: stay workdir-relative — they are run artifacts and land under the workdir, matching the checkpoint and .oxo-flow state.

This keeps a workflow self-contained: for a shared, read-only workflow (a fixed reference resource), point --workdir at the analysis directory instead:

# Workflow lives in a central, read-only location; data and results
# belong to the current analysis directory
oxo-flow run /opt/pipelines/wgs.oxoflow --workdir .

# Outputs, .oxo-flow checkpoint, and lock all land in the workdir;
# dry-run/clean/report/resume accept --workdir too, and resume re-uses
# the workdir recorded in the checkpoint

Parallel execution#

oxo-flow run pipeline.oxoflow -j 8

Keep going on failure#

oxo-flow run pipeline.oxoflow -j 4 -k

Retry failed jobs#

oxo-flow run pipeline.oxoflow -j 8 -r 2

Run specific targets#

oxo-flow run pipeline.oxoflow -t align -t sort

Target names may be prefixes: -t align matches every rule whose name starts with align (including expanded instance names like align_cohort_S1). The closure walks the instantiated DAG — it includes each target plus every upstream rule that survives the plan-time when gating, and skips when-gated variants entirely: a variant whose when is false under the effective config never enters the execution set, its upstream is not pulled in, and consumers of its (never-produced) outputs are pruned too. Naming a when-gated-false target explicitly reports it and aborts with "nothing to run" instead of planning a run that cannot produce output.

Set a per-job timeout#

# 3600 seconds, or equivalently:
oxo-flow run pipeline.oxoflow --timeout 1h

Limit resource usage#

# Use only 16 threads and 32GB memory
oxo-flow run pipeline.oxoflow --max-threads 16 --max-memory 32768

Cache environment setup#

# Cache environment setup state for faster subsequent runs
oxo-flow run pipeline.oxoflow --cache-dir .oxo-flow/cache

The cache would otherwise grow without bound, so after each run files untouched for 90 days are removed (the next run starts clean). The age limit is configurable per workflow — set cache_max_age_days = 0 to disable aging entirely, or any other value to change the window:

[config]
cache_max_age_days = 30   # prune after 30 days instead of the 90-day default

Skip environment setup (when environments are pre-built)#

oxo-flow run pipeline.oxoflow --skip-env-setup

Backend preflight: fail fast on a missing environment backend#

After DAG construction and before any rule executes, run validates every unique environment spec among the pending rules — once per spec (rule templates fan out one spec across N samples, so the spec set is tiny even for large cohorts). When a required backend binary is missing (conda/mamba/pixi/docker/singularity/modules), the run aborts before anything executes:

2 environment backend(s) required by pending rules are unavailable; no rules were run:
  - conda: conda is not installed or not in PATH — required by 3 pending rule(s), e.g. `bwa_align` (environment: envs/alignment.yaml); install conda and ensure it is on PATH
  - pixi: pixi is not installed or not in PATH — required by 1 pending rule(s), e.g. `report` (environment: pixi.toml); install pixi (https://pixi.sh) and ensure it is on PATH

Rules sharing one broken spec, and distinct specs failing with the identical backend-missing message, collapse into a single entry — one root cause per line, with the pending-rule count, an example rule, the primary spec, and a mode-aware remediation hint. Nothing ran, so the --json summary still reports the run as failed with zero counts (issue #142 H6). See Troubleshooting — Environment Issues for per-backend checks.

The remediation depends on why the spec failed (issue #848):

  • Manifest file not found (pixi declared by relative path): the spec is resolved against the workflow directory (the same base the executor uses, issue #427), and the hint says to create the manifest at the resolved path or correct the rule's pixi path — installing the backend would not help when the file itself is absent. Other backends keep the install advice (a missing conda YAML surfaces only at task execution).
  • Backend binary missing (the manifest exists, or the backend cannot be located for another reason): the hint is the per-backend install advice (install pixi (https://pixi.sh) and ensure it is on PATH, and so on).

This is the run-time safety net. validate and plan/dry-run preflight pixi manifests up front (error E019) so the failure surfaces before you commit to a run; run keeps the check because manifests can also be created between validate and run.

Pass config values#

Every key in the [config] section of the workflow (including declarative key = { default = …, required = …, … } entries) can be overridden from the CLI — no extra flags to declare:

# Direct flag form
oxo-flow run pipeline.oxoflow --database refs/nt

# Attached form
oxo-flow run pipeline.oxoflow --database=refs/nt

# Key=value form directly after the workflow file
oxo-flow run pipeline.oxoflow database=refs/nt threshold=1e-3

# Legacy --arg form (still supported)
oxo-flow run pipeline.oxoflow --arg database=refs/nt --arg threshold=1e-3

# Empty string = explicit derive-if-empty override
oxo-flow run pipeline.oxoflow gene_bed=
oxo-flow run pipeline.oxoflow --gene_bed ""

Empty values select the derive-if-empty branch

An empty value is accepted and overrides the config key with the empty string. Pipelines that auto-derive a value branch on [ -n "{config.gene_bed}" ] — passing gene_bed= forces that branch off (e.g. to re-derive from a new reference instead of reusing a declared default path). Only the key must be non-empty; =value alone is a syntax error.

Ordering: run flags before overrides

run flags (--json, --rerun, --samples, -j, …) must come before positional KEY=VALUE overrides — the override list is trailing, so a flag typed after it is reported with actionable guidance:

oxo-flow run pipeline.oxoflow min_quality=30 --json
'--json' is a command flag, not a config override.
  Command flags must come before KEY=VALUE overrides, e.g.:
  oxo-flow run <workflow.oxoflow> --json min_quality=30
  For a config key that itself starts with dashes, use --arg KEY=VALUE
  (also placed before positional overrides).

Execution profiles#

--profile <NAME> applies a reusable config supplement to a run. Profiles are plain TOML (or .oxoflow) files placed in the workflow's own profiles/ directory; .toml is tried before .oxoflow:

<workflow-dir>/profiles/<NAME>.toml     # preferred
<workflow-dir>/profiles/<NAME>.oxoflow

The name is just a filename — it carries no built-in meaning. What the profile contains decides its effect: a [config]/[defaults] block supplements config, a [cluster] block additionally routes execution to a scheduler (see Cluster submission). Naming a profile conda or slurm is a convention, not a mechanism — compute environments are declared per rule via [rules.environment] and are independent of profiles.

# profiles/cluster.toml
[config]
threads = "32"
memory_mb = "128000"
oxo-flow run pipeline.oxoflow --profile cluster

The profile's [config] table fills in keys the workflow does not set itself by default — values already present in the workflow are never overwritten. This lets one workflow carry per-environment fallback defaults (laptop, cluster, cloud). Declare profile_mode = "override" in the workflow's [workflow] section when a profile should instead replace workflow values (deep-merged: nested tables merge recursively, scalars and arrays are replaced) — the classic "cluster profile switches threads/memory" case:

[workflow]
name = "rnaseq"
profile_mode = "override"   # fill | override (default: fill)

If the named profile file does not exist, run (and dry-run) fail loudly with a nonzero exit and an error listing every available profile in the profiles/ directory — a typo'd --profile must never silently run with the workflow's own config. The same hard error applies when the profile file exists but its TOML fails to parse or merge.

Cluster submission#

A profile that also carries a [cluster] block makes run --profile submit to a scheduler instead of executing locally:

# profiles/slurm.toml
[cluster]
backend       = "slurm"      # slurm | pbs | sge | lsf — required
partition     = "compute"
account       = "lab01"
walltime      = "24h"        # a rule's own time_limit wins
max_submitted = 50           # jobs in flight at once
poll_interval = "30s"
extra_args    = ["--exclusive", "--constraint=haswell"]  # verbatim directives

[config]
reference = "/data/ref/hg38.fa"

Keys the profile omits fall back to the driver defaults: max_submitted 50, max_array_size 1001, poll_interval 5s. extra_args entries are emitted as scheduler directives verbatim and are not validated — a typo reaches the scheduler as written.

oxo-flow run pipeline.oxoflow --profile slurm

# one-off override of the profile's queue cap
oxo-flow run pipeline.oxoflow --profile slurm --max-submitted 100

Profiles are looked up in the workflow's own profiles/ directory — there is no user- or system-level profile location, so a site profile shared across a lab is copied into each workflow (or kept in the workflow repository) rather than installed once per machine.

Both conditions are required: a [cluster] block must be in effect and a profile must be named. A profile with only [config] keeps the local path, and a workflow carrying its own [cluster] block still runs locally until you pass --profile, so adding one never changes an existing run.

Rules become jobs after wildcard expansion, so a 100-sample scatter is 100 jobs, and dependencies are chained per instance. The run submits at most max_submitted jobs at a time, tracks them to completion, and writes the checkpoint exactly as a local run does — status, resume, and re-run invalidation behave identically, and a second run --profile slurm with nothing changed submits nothing.

Job arrays#

Same-template instances that are ready together and agree on every scheduler-visible directive (threads, memory, time limit, partition, account, GPU spec, environment, workdir) are grouped into one scheduler array automatically (issue #74):

[cluster]
backend        = "slurm"
max_array_size = 1001   # optional — the scheduler's MaxArraySize; larger
                        # scatter groups are chunked into several arrays
  • One submission instead of N sbatch/qsub calls — one queue entry that expands internally (SLURM --array=1-N, PBS -J, SGE -t, LSF -J "name[1-N]").
  • Array elements are tracked individually: per-instance job directories stay greppable (jobs/<instance>/job.sh + job.id + status.json), and index.json in the run directory maps each array index back to its instance, so the array never leaks into the human-facing layout.
  • Instances that differ on any directive fall out of the group and submit as single jobs — heterogeneity degrades gracefully.
  • Arrays are transport-level: the JobRecord set, checkpoint, and resume semantics are identical to per-job submission.

Element-wise aftercorr chaining is not part of this slice: downstream instances submit in ready batches as array elements finish, which is correct but not as queue-efficient as true element-wise chaining.

Partial module runs (--module)#

A composed workflow (clinical pipelines like clindet: several [[include]] modules chained together) can run just ONE module and the producers of its declared inputs:

# run only the germline module (rules/20_germline.oxoflow) of a composed
# clinical pipeline — the module name defaults to the include's file stem
oxo-flow run main.oxoflow --module 20_germline

# multiple modules + rule targets union together
oxo-flow run main.oxoflow --module 20_germline --module 30_vcf_norm -t report

The closure is: the module's rules + every host rule producing one of its declared concrete contract inputs; upstream DAG dependents come through the regular target machinery. Unknown module names fail with the known module list. dry-run --module previews the same set.

Run directory layout#

Each run leaves a directory you can navigate with ordinary tools:

.oxo-flow/runs/2026-08-17T14-30-05/   # `latest` symlinks to the newest
  events.jsonl                        # append-only submit/complete/fail log
  index.json                          # array index → instance name
  jobs/<rule>/job.sh                  # the exact script submitted
  jobs/<rule>/job.id                  # scheduler job id
  jobs/<rule>/status.json

Every events.jsonl line carries an RFC 3339 ts, so the log reconstructs a timeline rather than just an ordering:

{"ts":"2026-08-17T14:30:05.114Z","t":"SUBMITTED","rule":"align_S1","job":"4812345"}
{"ts":"2026-08-17T15:02:44.907Z","t":"COMPLETED","rule":"align_S1","job":"4812345"}

What a finished job records#

When a job leaves the queue, oxo-flow reads the scheduler's accounting store (sacct, qstat -x -f, qacct, bacct) and writes what it found into status.json:

{
  "state": "COMPLETED",
  "job_id": "4812345",
  "command": "bwa mem ref.fa data/S1.fq > aln/S1.bam",
  "submitted_at": "2026-08-17T14:30:05.114Z",
  "finished_at": "2026-08-17T15:02:44.907Z",
  "queue_wait_secs": 65,
  "exit_code": 0,
  "elapsed_secs": 1894,
  "max_rss_mb": 24680,
  "cpu_seconds": 14203
}

queue_wait_secs is submit-to-finish minus the scheduler's own elapsed time — the driver never observes the moment a job starts, so the split between waiting and working is only available by subtraction. The same elapsed_secs is what benchmarks.wall_time_secs records in the checkpoint, which therefore measures runtime rather than runtime plus queue wait, and max_rss_mb / cpu_seconds populate the same benchmark fields a local run fills from its own sampler.

How much of this appears depends on the site. Accounting is a deployment choice: a SLURM cluster without slurmdbd answers sacct with nothing, and LSF's bacct columns vary too much between versions to parse blind, so LSF records state only. Fields the store did not report are omitted rather than guessed — including exit_code, which stays absent for a failed job whose real code could not be read, instead of defaulting to 1.

Outputs, the working directory, and logs are assumed to live on storage shared between the submitting host and the compute nodes.

Not yet implemented. The driver polls until every job reaches a terminal state; it does not yet re-attach to jobs still in flight if the driver exits, and Ctrl-C does not yet cancel submitted jobs. Until then, use oxo-flow cluster status (no ids: lists your queued jobs) and oxo-flow cluster cancel <JOB_ID>... to inspect or clean up an interrupted run.

Cluster scheduler submission (SLURM, PBS, SGE, LSF) is configured separately — see the cluster command.

Execute a published bundle#

# Run from a bundle (extracts, verifies checksums, shows resource requirements,
# then prompts for confirmation — requires an interactive terminal on both
# stdin and stderr)
oxo-flow run --bundle pipeline-bundle.tar.zst -j 16

# Skip confirmation for CI/scripts (also required when stdin is redirected
# or with --json, which is always non-interactive)
oxo-flow run --bundle pipeline-bundle.tar.zst -j 16 --yes

# .tar.gz format also supported
oxo-flow run --bundle pipeline-bundle.tar.gz -j 16 --yes

# Pull from remote and execute in one step
oxo-flow pull gh:user/repo@v1 && oxo-flow run --bundle repo-bundle.tar.zst -j 16 --yes

Pilot runs and scale-up#

For a large cohort, run a fast pilot on a subset first; the checkpoint then skips the pilot samples when you scale up. --samples needs a workflow that has a sample dimension (sample_pattern, [[sample_groups]], sample_groups_file, or [[pairs]]): a workflow with none rejects the flag with --samples matched no samples in this workflow.

# Pilot: run the full pipeline on the first 2 samples only
oxo-flow run pipeline.oxoflow --samples first:2

# Or name the pilot samples explicitly (combines with first:N)
oxo-flow run pipeline.oxoflow --samples first:2,NA12899

# Scale up: no flag needed — completed samples are skipped automatically
oxo-flow run pipeline.oxoflow

# Preview the pilot plan without executing — with a checkpoint present,
# dry-run also predicts what will actually re-run (and what stays
# protected); see dry-run's "Checkpoint-Aware Rerun Preview"
oxo-flow dry-run pipeline.oxoflow --samples first:2

# Fix a config bug found by the pilot — only affected rules re-run
# automatically (see "Config changes" below); use --rerun to force everything
oxo-flow run pipeline.oxoflow --rerun

--samples filters every sample source (sample groups, sample_pattern discovery, and experiment/control pairs). --rerun re-executes the rules selected for this run even when outputs are up to date — combine it with --samples to re-run just the pilot subset while keeping the other samples' completed records.

After a --samples run, a pilot summary is printed: samples run, wall time, per-sample time, and a linear projection for the full cohort. Scientific-preflight findings are reported as distinct findings with their instance counts — one template rule appearing once per sample shows as 1 distinct finding (3 instance(s)), not 3 finding(s). When the workflow enables [ai], a plain-language pilot report (health assessment and scale-up advice) is appended automatically.

Incremental data arrival: --samples ready#

Sequencing centers deliver data in batches, so a cohort's fastq files trickle in over days. Instead of waiting for every sample, run the analysis as data arrives:

# Which samples can be processed right now, and what is still missing?
oxo-flow dry-run pipeline.oxoflow
#   Sample readiness: 87/100 complete, 13 waiting
#     ⏳ NA12891 (missing: data/NA12891_R2.fastq.gz), …

# Run only the samples whose entry inputs are complete
oxo-flow run pipeline.oxoflow --samples ready

# … more data arrives …
oxo-flow run pipeline.oxoflow --samples ready
#   Done: 5 succeeded, 87 skipped — the checkpoint skips completed samples

A sample is ready when every external input belonging to it exists — that is, every rule input (after wildcard and {config.x} expansion) that the workflow itself does not produce. Intermediate products are never checked (producing them is the DAG's job), and optional = true rules do not block readiness (the executor skips them when their inputs are absent). Relative paths resolve against the workflow file's directory (or --workdir) — the same place rules actually run from.

Semantics:

  • ready is a special --samples value, not a new syntax: it resolves to the names of ready samples and combines with first:N and explicit names as a union. Because it is reserved, a sample literally named ready cannot be selected by name.
  • With zero ready samples run aborts and lists the waiting samples; dry-run --samples ready instead reports the full cohort state.
  • For [[pairs]] workflows, a pair is kept only when both experiment and control inputs are complete; otherwise it is skipped with a note. Missing files that belong to no specific sample (shared references) are reported as workflow-level inputs.
  • The report is also available as JSON via dry-run --json (the samples block), and --samples ready works in test --run as well.

Background runs (--background)#

run --background (and resume --background) detaches execution into a background process and returns to the shell immediately:

# Foreground prints one line and exits 0; the run continues in the background
oxo-flow run pipeline.oxoflow --background
#   started in background (pid 48291) · log: .oxo-flow/logs/oxo-flow.log ·
#   monitor: oxo-flow status .oxo-flow/checkpoint.json · stop: kill 48291

# Detach a resumed run the same way
oxo-flow resume .oxo-flow/checkpoint.json --background

Mechanics:

  • The foreground process spawns a detached child that re-runs the same command with --background removed — every other flag (-j, --profile, --samples, --rerun, --log-file, …) passes through verbatim, so a background run behaves exactly like a foreground one: checkpoint, workdir lock, report snapshots, and resume semantics are unchanged.
  • The child's stdout and stderr are redirected to the run log (--log-file if given, otherwise .oxo-flow/logs/oxo-flow.log in the workdir). In this mode the tracing tee is skipped on purpose — the redirect already captures everything, and a second writer would duplicate every line (issue #194 A3).
  • The run log archives BOTH the structured tracing stream AND the user-facing progress narrative (Running: / ✓ / Done:) plus one JSON line per execution event (workflow_started, rule_started, rule_completed, workflow_completed) — the whole run is replayable from the single file (issue #194 B1/B3).
  • The child's pid is written to .oxo-flow/background.pid in the workdir (also shown in the summary) while the run is in flight; the file is removed when the child exits, so a stale pid never survives a finished run.
  • The child gets its own process group, so it survives terminal close and Ctrl-C at the shell — it keeps running until it finishes or is killed.
  • The foreground exits 0 once the child is spawned; it does not wait. If the spawn fails (for example because another run already holds the workdir lock), the foreground exits nonzero with the reason.

Which commands accept --background:

  • run and resume — the only long-lived terminal-bound commands. Cluster runs go through run --profile <name> and therefore detach the same way: run --profile slurm --background submits and polls the cluster jobs entirely inside the background process.
  • dry-run intentionally does not: a preview is fast and read-only — keep it in the foreground.
  • Other commands (pull, test, export, ai, …) are either quick or one-shot; run them before the background run and keep them foreground.

Monitoring a background run works exactly like a foreground one: poll oxo-flow status .oxo-flow/checkpoint.json (or status --timing), read the run log, and check the report snapshots in .oxo-flow/reports/. Stop a background run with kill <pid> (Unix) or taskkill /PID <pid> — the pid is in the summary line and (while the run is in flight) in .oxo-flow/background.pid.

Notes:

  • The global --json is rejected with --background: the foreground only launches the detached child, so there is no run summary to emit. Drop --json or run in the foreground to get the machine-readable summary.
  • Bundle runs need an explicit --workdir (the bundle's extracted directory is created per-process, so it cannot be predicted from the foreground invocation) and --yes — a detached child cannot prompt. --workdir is also what bundle discovery consults as a fallback: a sample_pattern/pairs_pattern that matches nothing inside the extraction is retried against the workdir (#751).
  • --background is a launcher flag, not a workflow setting: the child re-parses the remaining argv, so flag ordering rules (command flags before KEY=VALUE overrides) apply to the whole command line as usual.

Checkpointing and Resuming#

oxo-flow automatically persists execution state to a checkpoint file after every rule completion. This enables:

  1. Resuming failed runs: If a job fails or is interrupted, simply run oxo-flow run again. The engine will skip all rules already marked as completed and resume from the first pending task.
  2. State Inspection: Use the oxo-flow status command to view the progress of a run from its checkpoint file.

Metadata Directory#

By default, checkpoints are saved in a hidden .oxo-flow/ directory located in the same folder as the workflow file.

  • Filename: checkpoint.json (the name is always the same regardless of workflow name)

Rule run records in the checkpoint#

Each executed rule carries a rule_runs entry in checkpoint.json with the expanded command, the process exit code, a bounded stdout/stderr tail, the resolved report caption, and (since the cancellation-semantics change) the terminal status ("failed", "cancelled", "timed_out", …), the skip_reason when the rule did not run to completion, and the terminating signal when the process died to one. The three newer fields are optional and absent in checkpoints written by older engines — consumers must infer legacy records as before (failure when the rule is in failed_rules).

Timestamp semantics (issue #850): the engine stamps enqueued_at when an execution attempt is submitted and stamps started_at only after the rule has acquired its resources — the moment actual execution begins. Their difference is the time the rule spends waiting in the resource pool (same split as the cluster driver's queue_wait_secs), and all wall-time reporting (duration_ms, the report's Benchmarks Wall Time column, benchmarks.wall_time_secs) is measured from started_at, so a rule parked behind busy resources never inflates its reported runtime. These two timestamps live on the engine's in-memory job records and are not part of rule_runs entries; the checkpoint persists the resulting wall time via benchmarks.wall_time_secs, and cluster runs additionally persist the split per job as queue_wait_secs in jobs/<rule>/status.json (see What a finished job records).

The tails are truncated to the last 64 KiB per stream by default and prefixed with a … marker when cut. Set OXO_FLOW_OUTPUT_TAIL_BYTES to change the size (bytes; an invalid or zero value falls back to the default with a warning), e.g. OXO_FLOW_OUTPUT_TAIL_BYTES=10485760 (10 MiB) for error-heavy tools whose message scrolls past the tail. The in-RAM capture window widens to match when the setting exceeds its 1 MiB floor — budget roughly 2 × (1 MiB + tail) of RAM per running job. Tails are bytes-bounded, so a multi-byte UTF-8 cut lands on the nearest character boundary (rounding down); older checkpoints that only carry the legacy 2048-character tails remain readable (the new size applies from the next recorded run onward).

Declared log files vs captured tails#

A rule's declared log path is separate from this capture: the engine creates the log's parent directory, but the file only gains content when the rule's shell redirects into it (2> {log}) and the tool writes to stderr. Many bioinformatics tools self-log to their own location instead (e.g. a logs/<date>.txt file they manage themselves), leaving the declared log empty — an empty declared log never means no diagnostics were captured; the tails above are persisted in rule_runs regardless. When a rule succeeds, its declared log is empty (0-byte file or empty directory), and both captured tails are empty too, the run emits one Info-level hint that the tool likely self-logs elsewhere — quiet tools with output in any of the three stay silent (issue #832).

Workflow version in the checkpoint#

When the workflow file lives inside a git repository, oxo-flow records the repository's HEAD commit SHA in the checkpoint (workflow_git_sha) at run start. Every result set is therefore auditable to the exact workflow version that produced it: oxo-flow provenance verify prints the SHA, and report snapshots carry it through the checkpoint they embed.

The lookup is best-effort — running outside a git repository simply omits the field and never fails the run. Commit your workflow before running to get fully version-audited results. See Workflow Versioning for the full model.

Run logs#

Every run — and resume, which re-enters the same path — archives its own log under the workdir:

  • .oxo-flow/logs/oxo-flow.log is always the latest run; on each new run the previous log rotates to .oxo-flow/logs/oxo-flow.log.1, then .2, … up to .9 (the oldest backup is deleted). --log-file PATH overrides the location; the same rotation applies to PATH.
  • The log header names the exact workflow version that produced the record: timestamp, oxo-flow version, full command line, workflow path, workflow_name, workflow_version, git_sha (the repository HEAD recorded by the run — see Workflow Versioning), and workdir.
  • The engine's tracing stream (rule start/finish verdicts, warnings, errors) is written into the log alongside stderr. Logging is best-effort: if the file cannot be opened or written, the run continues with stderr-only logging and a warning. Dry-run remains stderr-only.

Report snapshots#

After every run (and resume) — unless --no-report-snapshot is given — oxo-flow writes a JSON report snapshot of the final checkpoint:

  • .oxo-flow/reports/report-<UTC timestamp>.json — the full report data model, so a run leaves a machine-readable report behind with no reporting step needed (a -N suffix is used when two snapshots land in the same second)
  • .oxo-flow/reports/index.json — a JSON array of {generated_at, workflow, checkpoint, report} entries, kept sorted by generated_at; the last entry is the newest snapshot

Snapshot failures are warnings — a reporting hiccup never fails a run. oxo-flow report with no -o writes auto-discovered reports into the same .oxo-flow/reports/ directory.

Config changes and precise invalidation#

The checkpoint records a snapshot of the effective config values and a structural fingerprint of every completed rule. On each run, oxo-flow compares them against the current workflow:

  • Changed config keys invalidate exactly the rules that reference them plus their DAG downstream; unrelated completed rules keep their checkpoint records. How a rule references a key decides the verdict:
  • A key interpolated into shell, script, input/output paths, envvars, or params ({config.<key>}) always invalidates — the new value bakes straight into the command or paths.
  • A key referenced only inside a when condition invalidates only when the condition's truth value actually flips under the new config. A toggle that leaves the gate true (or false) — e.g. (config.refine_bins || config.run_checkm) switching which term is true — reuses the completed rules instead of re-running hours-long chains whose inputs and commands are identical (issue #198). Flipping a gate in either direction still invalidates: the engine pre-marks completed rules without re-evaluating when, so only plan-time detection can retire a now-skipped producer or cascade a newly-activated one to its consumers.
  • Edited rule definitions (shell, inputs, outputs, environment, …) are caught by the per-rule fingerprint and invalidate the same way. A file-backed environment spec (pixi manifest, conda/mamba YAML, venv requirements) is fingerprinted by content, not just by the path the rule names — pixi specs also cover the sibling pixi.lock, since a lock re-pin changes exactly which packages the environment materializes. Editing any of these files in place invalidates the rules that reference them. Checkpoints written before the content tagging lack these digests, so the first run after that upgrade re-runs rules with file-backed environments once (the same safe-but-costly bootstrap as the config-tracking baseline below).
  • samples_list / samples_<group> are engine-injected and never trigger invalidation, so toggling --samples between runs stays cheap. This covers rules whose input list is baked from an injected key (expand_inputs over config.samples_list, a gather-rule pattern): their rule fingerprint and input manifest change with the selection, and oxo-flow treats that change as non-invalidating. When such a rule is skipped because of a selection change, run prints a warning naming the rule — its outputs still reflect the previous run's full sample set, and --rerun is the documented way to regenerate them under the new selection. The exclusion is narrow and verified: the rule's other definition fields must be byte-identical, and the current input file set must be exactly reproducible from the injected value with unchanged content — any genuine edit still invalidates. Checkpoints written by older binaries lack the input-excluded fingerprint, so the first subset run after an upgrade may re-run such a rule once (safe default).
# min_quality is interpolated by fastp_trim; only it and its downstream re-run.
# A flag toggled in a when-only rule whose gate stays true re-runs nothing —
# the run reports those rules as reused instead.
oxo-flow run pipeline.oxoflow min_quality=30
Config change:
  min_quality: 20 → 30
  when-condition unchanged, reused: bin_classify
  → invalidated 3 (1 directly affected), re-running 3/10 this run, skipping 7
  ⊝ index_reference (already completed)
  Running: fastp_trim
  …

When a when gate closes at execution time, the rule is reported under Skipped (when condition false) with a ⊘ marker (or ⊝ in the config-change reuse view) — a gated-off rule is a skip, never a failure (see status for the checkpoint-side display, including legacy checkpoints that carry stale failed_rules entries).

Rules changed by a transformed split set (transform.split.values_from) are detected through their baked input lists, so chunked rules re-combine correctly when the split values change.

Limitations (correctness-first, documented deliberately):

  • Changing threads/memory (performance knobs) does not invalidate results; changing anything else in a rule — including shell comments — does.
  • A checkpoint written before config tracking was introduced is adopted as a one-time baseline: the first post-upgrade run reuses everything and records the snapshot; changes made before that baseline cannot be detected. The same applies to the per-rule when verdicts behind gate-aware reuse: a pre-verdict checkpoint invalidates when-referencing rules once on the first changed-key run, records their verdicts, and reuses stably from then on.
  • Config values declared sensitive are stored in the snapshot as SHA-256 digests, never as plaintext.
  • Concurrent oxo-flow run invocations on the same workdir are prevented by an exclusive lock on .oxo-flow/lock (held for the whole run). The second invocation fails with a clear error instead of silently racing on checkpoint.json; the lock releases automatically when the first process exits (even on a crash). clean refuses to delete while a run is active unless --force is given.

Input changes and manifest invalidation#

The checkpoint also records an input manifest for every completed rule: the file set its inputs resolved to at completion time, with each file's path, size, and modification time — plus a content hash for files up to 64 MiB. On every run, oxo-flow re-resolves the inputs and compares:

  • Literal glob inputs (data/*.txt) detect added and removed files — a new file matching the glob rebuilds the rule (and its DAG downstream).
  • Directory inputs (e.g. input = ["results"] — a path that resolves to a directory) are listed recursively, so files added or changed anywhere inside invalidate the rule. Dir inputs with a pattern filter track only matching files.
  • Plain file inputs are content-addressed up to 64 MiB: a same-size rewrite is detected even when the mtime is preserved, and a mere touch no longer invalidates. Larger files compare size and mtime, closing the same stale-reuse hole for ordinary files.
  • expand_inputs-baked inputs whose producer is when-gated off under the active config are tolerated as absent: the manifest records the existing subset instead of erroring, so a no-op rerun does not invalidate the rule and cascade its producers (issue #757). If a later config change produces the missing files, the manifest then mismatches and the rule invalidates — the flip is still caught.

  • Missing external source inputs fail fast before any work starts: a required input that no rule in the DAG produces and whose concrete files are absent aborts the run before scheduling with workflow is missing required source input(s) — no rule produces them, naming each affected rule and pattern (exit non-zero). Previously a fully-completed checkpoint plus a deleted raw data file could exit 0 from stale checkpoint credit ("re-running 0 producer rule(s)"); that silent-success hole is closed. The gate never fires for generated intermediates (they have producers and are governed by the manifest cascade below), [[references]] outputs (built later in the same run), inputs discoverable through the workflow's sample_pattern directory, when-gated-off rules, zero-fan-out scatter rules, or optional rules.

# a new file appears in data/ after the last run
oxo-flow run pipeline.oxoflow
  ↻ input changes invalidated 2 rule(s): gather, report
  Running: gather
  Running: report
  …

Invalidated rules bypass the executor's freshness gate and re-execute even when their outputs exist and look recent; downstream rules that consumed their outputs rebuild in the same run.

Limitations (correctness-first, documented deliberately):

  • Detection is a hybrid policy: files up to 64 MiB are content-hashed (sha256), so a same-size rewrite is detected even when the mtime is preserved — and a mere touch no longer invalidates. Larger files (multi-gigabyte intermediates) keep the size+mtime policy, like make and Snakemake's default mode: hashing every input on every run would cost O(total input bytes), prohibitive for bioinformatics-scale files. A large file rewritten with identical size and a deliberately preserved mtime is not detected; use --rerun to force. Legacy checkpoints (entries written before hashing) keep comparing size+mtime instead of invalidating everything once.
  • Only one glob level is tracked per input pattern; brace-expansion patterns (*.{txt,log}) are not expanded by the manifest scanner, so their matched set is not tracked (the shell still expands them at runtime).
  • Symlinked directories are recorded as single entries, never traversed (cycle-safe), so changes inside a symlinked directory are tracked at the symlink's own metadata level.
  • Rules that clean up their inputs at the end of a successful run (transform.cleanup = true chunk consumers) are not manifest-tracked: their inputs are engine-managed intermediates governed by the upstream-cascade invalidation instead.
  • Optional rules with permanently-absent inputs (optional = true or optional = "any", issue #633): when a declared input cannot be resolved because its producer was never instantiated (an input_groups pattern that matched no files, or an endedness-filtered branch that never ran), the snapshot skips that entry and records a manifest for the inputs that do exist — the rule stays skipped/completed across runs instead of cascading a full re-run every invocation. The same grace applies during invalidation detection: a missing input with no producer in the DAG is absent by design for an optional rule, while a required rule with a missing input still invalidates (a deleted raw data file must re-run). A failed manifest snapshot is reported on stderr and in the run log (⚠ could not snapshot inputs of rule …) instead of being silently dropped.
  • A checkpoint written before input tracking was introduced adopts the current file set as a one-time baseline (the first post-upgrade run reuses everything); changes made before that baseline cannot be detected.

Temporary rules (temporary = true)#

Pipeline intermediates are usually the largest files in a run (per-sample BAMs, unsorted alignments) and often useless once every downstream rule has consumed them. Marking a rule temporary = true deletes its outputs after a fully successful run — once every dependent has completed — and records a tombstone in the checkpoint:

[[rules]]
name = "align"
output = ["aligned/{sample}.bam"]
temporary = true

This is checkpoint-aware, not a blind delete:

  • A plain re-run skips the rule like any completed rule — the output is not regenerated just because it is missing.
  • When a dependent actually needs the outputs again (its own outputs were deleted, its inputs changed, or the config changed), the producer is regenerated first (lazy cascade-up), then the dependent runs — the same order as the original run.
  • Failed runs keep the outputs (the deletion only happens when nothing failed), so debugging a broken run never has to re-fetch intermediates.
  • Leaf rules (no dependents) keep their outputs — there is no downstream work for them to enable.
  ↻ temporary outputs needed again — re-running 1 producer rule(s): trim_S1
  …
  ⊘ temporary outputs deleted for 'trim_S1' (1 file(s), regenerated on demand)

Dry-run predicts exactly when a tombstoned rule will regenerate ([rerun: upstream of X]) — see Checkpoint-aware rerun preview.

Disk-pressure monitoring ([engine])#

Large pipelines can fill a disk mid-run. Declaring an [engine] section in the workflow makes the engine measure the working directory's free space at startup, at every scheduling round, and every 60 s while work is parked:

[engine]
min_free_disk = "50G"        # floor — entering Hold below it
reclaim_free_disk = "100G"   # optional reclaim trigger; defaults to min_free_disk

The response ladder:

  • Reclaim (min ≤ free < reclaim): completed rules' temporary outputs are tombstoned and unlinked mid-run — the same checkpoint-aware reclaim as temporary = true cleanup, but while the run is still going. Only when every dependent is completed and its own outputs still exist (a dependent whose outputs are missing will re-run and needs its inputs). The narration line is ⊘ reclaimed temporary outputs of '<rule>' ….
  • Hold (free < min): new rule dispatches are withheld; in-flight rules continue. If nothing is in flight and nothing is reclaimable, the run parks and re-measures every 60 s.
  • Abort (Hold persists ~2 parked rounds ≈ 2 min): the run checkpoints and fails. Free space manually (delete intermediates or oxo-flow clean), then oxo-flow resume — or lower min_free_disk.

Protected outputs (protected_output) are never reclaimed, and the final (leaf) rule's outputs always survive. See [engine] in the workflow format reference for the full field table and edge cases.

Tuning knobs (both fall back to their defaults when unset, invalid, or 0, so a typo can never disable the abort):

  • OXO_FLOW_DISK_TICK_SECS — seconds between measurements while parked (default 60).
  • OXO_FLOW_DISK_HOLD_ROUNDS — grace park rounds under Hold before the run aborts (default 2).

Known boundaries:

  • Mid-run reclaim of fan-out producers (output_pattern rules with per-instance values) does not fire — the event loop tracks reclaim candidates by rule name and holds no instance bindings, so those intermediates are only cleaned by the post-run sweep (instance-level, issue #844) or the next run. Fail-safe: no reclaim rather than a wrong one.
  • Free space is measured at the working directory's filesystem, which is volume-level (statvfs): thresholds apply to the whole volume, not to a quota.
  • Keep scratch-heavy tools off the measured volume only if they are not cleaned by their rules; outputs declared by rules are always accounted for by the ladder.

Forcing Execution#

To bypass checkpoints and re-execute rules that have already completed, use oxo-flow clean to remove outputs and the checkpoint file before running, or --rerun to re-execute this run's rules while keeping checkpoint records for rules outside the run.


Checkpoint re-entry (checkpoint = true)#

A checkpoint rule discovers new values at runtime: after it completes, the engine reads its checkpoint_manifest, merges the declared samples, and executes the newly created rule instances in the same run. The checkpoint records each re-entry (reentries), so a resume replays it and reconstructs the same plan; invalidating the checkpoint rule revokes its samples until it re-runs. A missing or unparsable manifest fails the checkpoint rule. See Workflow Format for the config surface.

Output#

oxo-flow v0.23.2 — Rust-native bioinformatics pipeline engine
DAG: 5 rules in execution order
  1. fastqc
  2. trim_reads
  3. bwa_align
  4. sort_bam
  5. call_variants
⠋ [00:15] [████████████░░░░░░░░] 3/5 ETA:0:00:42 (executing sort_bam)
  ✓ fastqc (2.1s)
  ✓ trim_reads (15.0s)
  ✓ bwa_align (42.3s)
  ✓ sort_bam (10.2s)
  ✓ call_variants (8.4s)

Done: 5 succeeded, 0 skipped, 0 failed
✓ 5 output files verified (118.3MB total)

A progress bar shows execution progress with:

  • Elapsed time
  • Current position / total rules
  • Estimated time remaining (ETA), plus the rule currently executing
  • The final summary line reports succeeded / skipped / failed counts

Notes#

  • The workflow file is optional; if not specified, auto-discovery searches for main.oxoflow first, then any *.oxoflow file alphabetically
  • If no .oxoflow file is found, an error message suggests running oxo-flow init to create one
  • The DAG is built and validated before any rules execute
  • Rules are executed in topological order; independent rules may run in parallel up to the -j limit
  • If --keep-going is not set, execution stops at the first failure
  • --keep-going changes scheduling, never the verdict: it runs the remaining rules past a failure, but the run still exits non-zero when a required rule failed — scripts and the web UI (which classifies by exit code) see the truth. Only required = false failures leave the exit code at 0.
  • The --retry flag re-runs failed jobs up to N times before marking them as failed
  • A timeout of 0 means no timeout
  • Resource constraints (threads, memory) in rules are checked against available resources before execution
  • Pre-flight budget check: When --max-threads or --max-memory is explicitly set, the engine checks all rules before execution starts. Any rule whose requirements exceed the budget is reported immediately (fast-fail), preventing mid-pipeline failures. An explicit budget is a hard limit.
  • Backend preflight: before any rule executes, every unique environment spec of the pending rules is validated once — a missing backend (conda, mamba, pixi, docker, singularity, modules) aborts the run with one collapsed entry per root cause and "no rules were run" (see Backend preflight).
  • Auto-detected capacity is soft: Rules whose declared requests exceed the machine's detected threads/memory are not rejected — the declared request is the tool's upper bound (often an upstream HPC label), not a scheduling requirement. The engine warns and clamps the pool reservation, so an over-capacity rule runs alone (serialized) instead of blocking the workflow.
  • Deadlock detection: If pending rules remain while nothing is running and none can become ready (typically an upstream failure), the engine reports Deadlock detected: N rules stuck with the stuck rule names. Resource waits cannot deadlock: over-capacity requests are clamped, and explicit budget violations fail fast before any rule runs.
  • Target-aware execution: The -t flag supports prefix matching — -t al matches all rules whose names start with "al". Use this to run a subset of the workflow, similar to make <target>. Only the named targets and their transitive upstream dependencies are executed; downstream rules are excluded.
  • Default targets: rules marked target = true are what run/dry-run build when neither -t nor --module selects anything — the marked rules and their transitive producers. With no marked rule the full DAG runs (backward compatible); an explicit -t/--module always wins.
  • Abort cancellation semantics: with fail-fast aborts, rules that were merely killed by the abort (signal death, no exit code) are recorded as cancelled, not failed — they are excluded from failed_rules and the failure counts, and narrated like skips. A sibling that failed on its own merit (exit code present) keeps its failure record. The checkpoint's per-rule rule_runs entries distinguish the two via their status/skip_reason/signal fields.
  • Setting --max-threads 0 or --max-memory 0 auto-detects system resources
  • Environment setup is performed automatically before first use of each environment (conda, pixi, docker, singularity, venv)
  • Use --skip-env-setup when environments are pre-built to avoid redundant setup
  • Use --cache-dir to persist environment setup state across runs for faster startup