Workflow Format#
The .oxoflow file format is oxo-flow's TOML-based workflow definition language. This page provides the complete specification, design philosophy, and syntax rules.
Design Principles#
The .oxoflow format is built on four core principles:
- Declarative over Imperative — Define what should happen (inputs, outputs, tools), not how to orchestrate it. The engine handles the execution logic.
- Explicit is better than Implicit — Every dependency and environment should be clearly visible. No hidden global state.
- Composition over Inheritance — Reuse logic through modular
includedirectives and rule templates rather than complex inheritance hierarchies. - Traceability by Default — The format structure directly supports generating provenance and audit trails (workflow checksums, per-rule input manifests, output hashes).
TOML Primer#
oxo-flow uses the TOML (Tom's Obvious, Minimal Language) format. If you are new to TOML, here are the three essential concepts used in .oxoflow files:
- Key-Value Pairs:
key = "value". Strings must be in quotes. - Tables:
[name]defines a section (an object/map). - Arrays of Tables:
[[name]]defines a list of sections. In oxo-flow, rules are defined using double brackets because a workflow contains multiple rules.
For more details, see the Official TOML Specification.
File Extension#
Workflow files must use the .oxoflow extension (e.g., qc_pipeline.oxoflow).
Top-level Structure#
reference_dir = "..." # Optional: base directory for auto-derived reference paths
[workflow] # Required: metadata
[config] # Optional: user variables (plain or declarative form)
[defaults] # Optional: rule defaults
[report] # Optional: report configuration
[[include]] # Optional: include external workflow files
[[references]] # Optional: auto-built indexes and reference data
[[rules]] # Required: one or more rules
[[pairs]] # Optional: experiment-control pairs (WC-01)
[[sample_groups]] # Optional: multi-sample groups (WC-02)
[[values]] # Optional: named value lists for parameter wildcards
[metadata] # Optional: per-sample metadata columns ({meta.<column>})
[resource_budget] # Optional: resource limits (enforced on cluster submissions; local runs use `--max-threads`/`--max-memory`)
[env_groups] # Optional: named reusable environment specs
[resource_groups] # Optional: shared resource pools (API limits, DB connections)
[wildcard_constraints] # Optional: regex patterns to constrain wildcard values
[cluster] # Optional: HPC cluster profile (SLURM, PBS, SGE, LSF)
[[reference_db]] # Optional: tracked reference database versions
[citation] # Optional: citation metadata (DOI, authors, etc.)
[plugins] # Optional: plugin configuration
[webhook] # Optional: workflow-level webhook notifications (issue #227)
[engine] # Optional: runtime disk-pressure guard (issue #843)
reference_dir is a bare top-level key, so it must appear before the first
[table] header; it is also accepted inside [config]. See
Auto-Derivation from reference_dir.
[[include]] — Modular Workflow Composition#
Include external workflow files to enable modular, reusable workflow design:
| Field | Type | Required | Description |
|---|---|---|---|
path |
String | Yes | Path to the included .oxoflow file |
namespace |
String | No | Optional namespace prefix for included rule names |
Namespace Behavior#
When a namespace is specified:
- All rule names from the included file are prefixed:
namespace::rule_name - Internal
depends_onreferences within the included file are automatically prefixed - External
depends_onreferences (to rules outside the included file) remain unchanged
Example:
# qc.oxoflow
[[rules]]
name = "fastqc"
input = ["{sample}.fastq.gz"]
output = ["qc/{sample}_fastqc.html"]
shell = "fastqc {input}"
[[rules]]
name = "trim"
input = ["{sample}.fastq.gz"]
output = ["trimmed/{sample}.fastq.gz"]
depends_on = ["fastqc"] # Internal reference - will become "qc::fastqc"
shell = "fastp {input} -o {output}"
# main.oxoflow
[[include]]
path = "qc.oxoflow"
namespace = "qc"
[[rules]]
name = "align"
input = ["trimmed/{sample}.fastq.gz"]
depends_on = ["qc::trim"] # Reference to included rule with namespace
shell = "bwa mem ref.fa {input} > aligned/{sample}.bam"
Resulting rules: qc::fastqc, qc::trim, align
Interface Contracts#
An [[include]] may declare a typed interface — what the host must wire
in, what the module exposes, and config defaults it reads (issue #112).
All checks run at validation time; execution semantics are unchanged,
and includes without a contract behave exactly as before.
# qc.oxoflow
[[rules]]
name = "fastqc"
input = ["raw/{sample}.fq.gz"] # the module's NEED
output = ["qc/{sample}_fastqc.html"]
shell = "fastqc {input} -o {output}"
[[rules]]
name = "internal"
input = ["qc/{sample}_fastqc.html"]
output = ["qc/{sample}.bin"] # internal — not declared below
shell = "true"
# main.oxoflow
[[include]]
path = "qc.oxoflow"
namespace = "qc"
inputs = ["raw/{sample}.fq.gz"] # host MUST produce these
outputs = ["qc/{sample}_fastqc.html"] # module EXPOSES these
params = { threads = "4" } # config defaults (host values win)
[[rules]]
name = "download"
output = ["raw/{sample}.fq.gz"] # wires the declared input
shell = "true"
| Field | Type | Description |
|---|---|---|
inputs |
Array of String | File patterns the host must wire into the module. A concrete declared input (no {-wildcard) that no rule produces fails validation with the gap named; wildcarded declared inputs are not statically checkable (a {sample} input forms no DAG edge — the placeholder only resolves at run time) — the host is responsible for wiring them |
outputs |
Array of String | Files the module exposes to the host. A declared output not produced by a module rule fails validation; a host rule reading a module-internal file that is not declared here produces an encapsulation warning. Wildcarded host inputs are matched structurally (identical literal prefix and suffix, e.g. qc/{sample}.tmp), so patterns that could address the same files warn too; patterns that resolve to different literals are not statically checkable |
params |
Table | Defaults for config keys the module reads, filled in profile-style (or_insert — explicit host values win). The module's own [config] entries fill any remaining gaps: a module declaring [config] trim_quality = "20" keeps that default when included without params, while any host value (its own [config] table or [[include]] params) still wins. A key referenced but defined nowhere remains an E005 error |
name |
String | Module identity for partial runs (run --module <name>). Defaults to the included file's stem (rules/20_germline.oxoflow → 20_germline), so existing composed workflows are addressable without changes |
repo + ref |
String pair | Pin the included path to a git repository at a tag/branch/commit: repo = "https://github.com/org/oxo-flow-modules", ref = "v0.23.2", path = "rules/qc.oxoflow". Any git URL works (https/ssh/file://); github.com clones fall back to China mirrors. Checkouts are cached under ~/.cache/oxo-flow/modules/<repo>@<ref>@<hash> (the hash disambiguates confusable repo/ref pairs; OXO_FLOW_MODULE_CACHE overrides the cache root) — versioned modules, reproducible composition |
[workflow] — Metadata#
[workflow]
name = "my-pipeline"
version = "1.0.0"
description = "A short description"
author = "Your Name"
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
name |
String | Yes | — | Pipeline name |
version |
String | No | "0.1.0" |
Semantic version |
description |
String | No | — | Human-readable description |
author |
String | No | — | Author name or email |
interpreter_map |
Table | No | {} |
Custom interpreter mapping for script extensions |
genome_build |
String | No | — | Genome reference build identifier (e.g., "GRCh38", "hg38") |
min_version |
String | No | — | Minimum engine version required to load this workflow: an older engine refuses it at parse time (compares major.minor.patch; "0.21" equals "0.23.2") |
format_version |
String | No | — | Format specification version for compatibility |
pairs_file |
String | No | — | External TSV/CSV/JSON file defining experiment-control pairs |
sample_groups_file |
String | No | — | External TSV/CSV/JSON file defining sample groups |
metadata_file |
String | No | — | External TSV/CSV/JSON file of per-sample metadata columns, addressed as {meta.<column>} (see Sample Metadata) |
pairs_pattern |
String | No | — | File glob pattern for auto-discovering pairs (e.g., "aligned/{pair_id}/{exp}_vs_{ctrl}.bam") |
sample_pattern |
String | No | — | File glob pattern for auto-discovering samples (e.g., "raw/{sample}_R1.fastq.gz") |
on_complete |
String | No | — | Shell command run once after the whole workflow completes successfully (no failed rules). Rendered with {config.*} and the counters {succeeded}/{failed}/{skipped} plus {workdir}; runs in the workflow root with the workflow shell_prelude. Best-effort: a failing hook warns, never changes the run status. |
on_error |
String | No | — | Shell command run once after the workflow reaches a terminal state with at least one failed rule. Same rendering and best-effort semantics as on_complete. |
Workflow-Level Terminal Hooks#
on_complete / on_error mirror the rule-level on_success / on_failure
hooks at the whole-workflow level — the nf-core workflow.onComplete /
snakemake onerror pattern (completion emails, cleanup, notifications):
[workflow]
name = "my-pipeline"
version = "1.0.0"
on_complete = "notify 'run done: {succeeded} succeeded, {failed} failed' --tag {config.project}"
on_error = "notify 'run FAILED: {failed} failed' --tag {config.project} && tar czf failed-run.tar.gz {workdir}/logs"
Both hooks run exactly once per oxo-flow run (the CLI) that reaches a
terminal state (including --keep-going runs); on_complete fires only
when no rule failed, on_error otherwise. The web UI executor has its
own run loop and does not fire these hooks yet. Placeholders: every
[config] key is available as {config.<key>}, plus {succeeded},
{failed}, {skipped}, and {workdir}. Unknown keys stay literal.
Custom Interpreters (interpreter_map)#
By default, oxo-flow auto-detects interpreters based on file extensions:
.py→python.py3→python3.R,.r→Rscript.sh,.bash→bash.jl→julia.pl→perl.rb→ruby.qmd,.Rmd,.rmd→quarto render.ipynb→jupyter nbconvert --to notebook --execute.smk→snakemake.nextflow→nextflow run.wdl→miniwdl run
You can override or extend this mapping in the [workflow] section:
[workflow]
name = "custom-interpreters"
[workflow.interpreter_map]
".m" = "octave"
".sas" = "sas"
".py" = "/opt/conda/bin/python" # Override default
This mapping applies only to the script field.
[config] — Configuration Variables#
User-defined key-value pairs accessible in rules as {config.<key>}. Every
config key can be overridden from the CLI.
Two forms are supported:
[config]
reference = "/data/ref/hg38.fa" # plain value (implicit default)
samples_dir = "raw_data"
results_dir = "results"
min_quality = "30"
Values are TOML strings, integers, booleans, or arrays. String interpolation
in rules uses {config.key} syntax; array values render as space-joined
lists in the shell.
Several keys are engine-injected:
config.samples_list(all sample names, comma-joined) andconfig.pairs_list(all pair ids, comma-joined) — use them inexpand_inputsto gather across samples or pairs without maintaining a parallel list by hand. A user-declaredconfig.samples_listis union-merged with the discovered samples, and the two lists are also rewritten by--samplesfiltering.config.samples_<group>(one per[[sample_groups]]entry).config.<reference name>(each[[references]]entry'soutput) and, withreference_dir, the ten derived paths (reference_fasta,gene_annotation,bwa_index,bwamem2_index,bowtie2_index,star_index,hisat2_index,minimap2_index,gatk_dict,samtools_faidx). Explicit[config]entries with the same names are never overwritten.
Declarative Form (inline table)#
For user-facing parameters, use the declarative inline-table form. Declared
keys become automatic CLI flags (--<key> <value>), can be required, and
carry metadata:
[config]
database = { required = true, help = "Path to the BLAST database" }
threshold = { default = "1e-5", help = "E-value cutoff" }
mode = { default = "dna", choices = ["dna", "rna"], type = "string" }
samples = { required = true, type = "path", must_exist = true }
Fields#
| Field | Type | Description |
|---|---|---|
default |
String | Default value when not provided via CLI |
required |
Bool | If true, the value must be provided at runtime (default: false) |
help |
String | Human-readable description shown in error messages |
sensitive |
Bool | Mask the value as *** in captured rule output, recorded commands, checkpoint/report snapshots, and error summaries — before they reach the web UI or AI recovery (default: false). Every non-empty value is masked; values of 4+ characters also mask their base64, percent-encoded, and JSON-escaped forms |
type |
String | Expected value type for validation: "string", "int", "float", "bool", "path" |
choices |
Array of String | Allowed values (requires type = "string") |
range |
String | Numeric range "min..max" (requires type = "int" or "float") |
must_exist |
Bool | Path must exist on disk (requires type = "path") |
Usage in Rules#
Values are accessible as {config.<name>} in shell commands, input/output
paths, and when conditions:
[[rules]]
name = "blast"
shell = "blastn -db {config.database} -evalue {config.threshold} -query {input} -out {output}"
CLI#
# Provide a required value (direct flag form)
oxo-flow run pipeline.oxoflow --database refs/nt
# Attached form
oxo-flow run pipeline.oxoflow --database=refs/nt
# Key=value form directly after the workflow file
oxo-flow run pipeline.oxoflow database=refs/nt
# Override several values
oxo-flow run pipeline.oxoflow --database refs/nt --threshold 1e-3
# Legacy --arg form (still supported)
oxo-flow run pipeline.oxoflow --arg database=refs/nt --arg threshold=1e-3
Precedence#
CLI override > declared default > error if required and unset.
An override key that differs from a declared key only by hyphens vs underscores
resolves onto the declared spelling: with threads_max = { default = "4" } in
[config], --threads-max=8 sets threads_max (and the --threads-max 8
space form is accepted too). This matters most for sensitive keys — masking
and the checkpoint snapshot are keyed on the declared spelling — but it applies
to every declared key. Keys matching nothing declared are untouched.
[[references]] — Auto-Built Indexes & Reference Data#
[[references]] is for external, pre-staged reference artifacts: the
engine builds them when missing, before the rule DAG runs. In-workflow
derivations (e.g. a STAR index produced by an upstream rule) must flow
through normal rule output/input edges instead — a reference entry
cannot consume rule outputs because references are built before the DAG
exists.
Declare reference artifacts (indexes, data files) that the engine auto-builds
when missing. Each [[references]] entry specifies a source, output, and build
command. The engine tracks built state in .oxo-flow/reference-checkpoint.json
and never rebuilds unnecessarily: each entry stores a fingerprint of the
definition plus the config values its build command references. The source
file's content state joins the fingerprint: size and mtime for every
source, plus a content hash for files up to 64 MiB — so for those files even
a timestamp-preserving same-path replacement (cp -p, rsync -t, git
checkout) triggers a rebuild; larger sources are guarded by size + mtime.
Editing the build/source/output, changing a referenced config value, or
touching the source file also triggers a rebuild, and rules that consume the
artifact through declared input paths (plus their downstream) are
invalidated so their outputs are regenerated.
When reference_dir is set, eight standard indexes are auto-derived without
explicit [[references]] blocks: samtools faidx, BWA, BWA-MEM2, Bowtie2,
minimap2, STAR, HISAT2, and the GATK sequence dictionary (see the
auto-derivation table below).
Use --skip-ref-build to skip automatic reference building.
Syntax#
[[references]]
name = "bwa_index"
source = "{reference_dir}/genome.fa"
output = "{reference_dir}/bwa/genome.fa"
build = "mkdir -p {reference_dir}/bwa && bwa index -p {output} {source}"
threads = 8
memory = "8G"
description = "BWA index for alignment"
[references.environment]
conda = "envs/bwa.yaml"
Fields#
| Field | Type | Required | Description |
|---|---|---|---|
name |
String | Yes | Unique name (used for checkpoint tracking) |
source |
String | No | Source file for freshness checks |
output |
String | Yes | Path to the built artifact |
build |
String | Yes | Shell command to produce output from source |
threads |
Integer | No | CPU threads for the build command |
memory |
String | No | Memory limit (e.g., "64G") |
description |
String | No | Human-readable description |
environment |
Table | No | Environment providing the build tool — same spec as [rules.environment] (conda/mamba/docker/…). Without it the build runs in the bare system shell, so builds that need workflow tools (bowtie2-build, STAR --runMode genomeGenerate, …) must declare one. The environment is created on first use and shared with rules through the same cache. |
Built-in Builder Templates#
Instead of a handwritten build shell, name one of the built-in
templates — the engine expands it into a canonical command (any string
containing whitespace or shell syntax is treated as a handwritten
command and passes through unchanged):
| Template | Produces |
|---|---|
samtools_faidx |
.fai index |
picard_dict |
.dict sequence dictionary |
bwa_index |
BWA's five index files (.amb/.ann/.bwt/.pac/.sa; the prefix is derived from output) |
star_index |
STAR genome directory (output = --genomeDir) |
bismark_index |
Bismark Bisulfite_Genome directory (source = FASTA directory) |
[[references]]
name = "genome_bwa"
source = "refs/genome.fa"
output = "refs/genome.fa.bwt"
build = "bwa_index"
threads = 8
Unknown template names, missing outputs, and duplicate reference names
are rejected at validate. Additionally, every reference injects a
keyed config value — config.<name> = its output — so rules consume
the artifact as {config.genome_bwa} without repeating paths.
Auto-Derivation from reference_dir#
When [config] contains reference_dir and no explicit [[references]] blocks
are declared, the engine automatically derives eight standard references:
| Name | Output |
|---|---|
samtools_faidx |
{reference_dir}/genome.fa.fai |
bwa_index |
{reference_dir}/bwa/genome.fa.bwt |
bwamem2_index |
{reference_dir}/bwamem2/genome.fa.0123 |
bowtie2_index |
{reference_dir}/bowtie2/genome.fa.1.bt2 |
minimap2_index |
{reference_dir}/genome.fa.mmi |
star_index |
{reference_dir}/star/SAindex |
hisat2_index |
{reference_dir}/hisat2/genome.fa.1.ht2 |
gatk_dict |
{reference_dir}/genome.dict |
reference_dir also auto-derives config defaults such as
reference_fasta ({reference_dir}/genome.fa) and gene_annotation
({reference_dir}/genes.gtf). Users can override any auto-derived reference
by declaring it explicitly.
CLI#
# Auto-build missing references
oxo-flow run pipeline.oxoflow -j 16
# Skip reference building (assume pre-built)
oxo-flow run pipeline.oxoflow --skip-ref-build
[defaults] — Default Settings#
Applied to all rules unless explicitly overridden:
[defaults]
threads = 4
memory = "8G"
environment = { conda = "envs/base.yaml" }
shell_prelude = "set -euo pipefail"
| Field | Type | Description |
|---|---|---|
threads |
Integer | Default CPU thread count |
memory |
String | Default memory allocation |
environment |
Table | Default environment specification |
shell_prelude |
String | Shell text prepended to every rule command, hook, reference build, and cluster job script on its own line (issue #92) — e.g. set -euo pipefail for fail-fast shells. Opt-in: empty by default, so existing workflows keep their exact command text. Applies inside environment wrappers (containers, conda) on every execution path. pipefail requires bash (the engine falls back to sh only when bash is absent) |
Conda/mamba environment naming#
For file-backed env specs the conda/mamba env name is derived as
<stem-or-yaml-name>-<hash8>, where hash8 is the first 8 hex chars
of the SHA-256 of the YAML content (issue #159). Relative env file
specs (conda, mamba, pixi, venv_requirements) resolve against
the workflow directory — the single base for relative paths in the
workflow file (issue #427); with --workdir the engine anchors them
automatically, so rule shells running from the workdir still find the
specs. venv (a directory the backend creates in the workdir) and
conda_prefix/mamba_prefix (an environment store outside the
workflow tree) stay relative to the workdir. Two workflows that
ship different YAMLs under the same name therefore build into distinct
envs instead of silently sharing one prefix; identical content keeps
the same name, so specs still deduplicate. Non-file specs (plain names
or inline strings) keep their plain name. Note: the first run after
this convention shipped rebuilds envs once under the new names —
remove the old plain-named envs with conda env remove -n <old> if
disk space matters.
[report] — Report Configuration#
Configure report generation behavior. Reports are built by a pluggable section system — each section is produced by a ReportSectionGenerator. Use sections to control which generators run.
[report]
template = "report.html"
sections = ["universal", "workflow-info", "commands", "clinical-compliance"]
| Field | Type | Description |
|---|---|---|
template |
String | Report template: the built-in name "report.html", or a template file path (workflow directory first, then cwd). Applies to HTML output only; a render failure warns and falls back to the default renderer |
format |
Array | Parsed but not supported yet — setting it makes report warn (or fail under --strict); select the output format with -f instead |
sections |
Array | Report sections to include. If empty (or omitted), all applicable generators run. Available built-in IDs: universal, execution-status, failure-diagnosis, clinical-compliance, workflow-info, commands, file-manifest, environment, metrics, sample-matrix, provenance, task-summary, software-versions, rule-captions, aggregate-metrics — 15 in total (run oxo-flow report --list-sections for the live list) |
How Sections Work#
Each section ID maps to a registered ReportSectionGenerator:
| Section ID | Description | When Active |
|---|---|---|
universal |
Dashboard with QC indicators and task counts | Always |
workflow-info |
Name, version, author, config, sample/pair counts | Always |
commands |
Expanded shell commands for every rule | Always |
file-manifest |
Input and output file listings | Always |
environment |
Available backends and oxo-flow version | Always |
clinical-compliance |
ACMG/AMP classification, audit trail, biomarkers | Always |
execution-status |
Per-rule execution status and benchmark metrics | Only with checkpoint |
software-versions |
Declared software per rule, nf-core-style: docker image:tag, module tool/version, env files with sha256 hashes — static declaration, runtime package versions deliberately not recorded (see Reporting) |
When any rule declares an environment |
rule-captions |
Per-rule report annotations rendered as markdown — see Rule Captions |
When any rule declares report |
aggregate-metrics |
MultiQC-style sample × metric matrix across tools, plus *_mqc.json custom content |
When metric or custom-content files are found |
The domain (DNA-seq, RNA-seq, epigenomics, or generic) is auto-detected from tool names in the workflow. Custom generators can be registered programmatically.
Rule Captions#
A rule can carry a report annotation — caption text rendered in the report's rule-captions section (the Snakemake report() equivalent). Two TOML forms:
[[rules]]
name = "qc"
shell = "fastp -i {sample}.fq"
# Inline caption (markdown allowed):
report = """
# QC summary
Reads trimmed with fastp; per-sample Q30 below.
"""
[[rules]]
name = "align"
shell = "bwa mem -t {threads} {input} > {output}"
# Structured form: a markdown/text file relative to the workdir,
# and/or an inline caption (at least one of the two):
report = { file = "notes/alignment.md", caption = "Alignment QC" }
At render time each executed rule instance contributes one markdown subsection, ordered by rule declaration order:
- The checkpoint's execution-time caption wins — it is what the run actually read (
report.filecontent is captured when the rule completes, so a later edit to the file does not silently change a finished report). - Without an execution record (dry run, or a checkpoint from before this feature), the declared annotation is resolved at report time: inline caption first, then
report.filerelative to the workdir. - A missing or unreadable
report.filenever fails the report — the rule simply contributes no caption.
Captions are text/markdown only; embedding figures is not supported in v1.
[[rules]] — Rule Definitions#
Each [[rules]] entry defines a pipeline step. The double brackets indicate a TOML array of tables.
Basic example#
[[rules]]
name = "align"
input = ["{sample}_R1.fastq.gz", "{sample}_R2.fastq.gz"]
output = ["aligned/{sample}.bam"]
environment = { conda = "envs/alignment.yaml" }
shell = "bwa mem -t {threads} {config.reference} {input} | samtools sort -o {output}"
[rules.resources]
threads = 16
memory = "32G"
All fields#
| Field | Type | Required | Description |
|---|---|---|---|
name |
String | Yes | Unique rule identifier |
input |
Array of strings | No | Input file paths. Optional (defaults to empty) — a rule with no input is legal (e.g. a pure producer or a setup step; pair it with depends_on when ordering matters) |
output |
Array of strings | No | Output file paths. Optional (defaults to empty) — a shell-only rule with no declared outputs validates and runs; pair with depends_on for ordering |
output_pattern |
String | No | Runtime-discovered output pattern: outputs enumerated by a filesystem scan after the rule completes; downstream consumers of its wildcard instantiate per discovered value. A [[values]] wildcard in the pattern is a plan-time fan-out trigger instead (issue #296). Mutually exclusive with output/transform/input_groups — see Runtime-Discovered Outputs |
shell |
String | No | Shell command to execute |
script |
String | No | Script file path (auto-detects interpreter) |
description |
String | No | Human-readable description of what this rule does |
report |
String or Table | No | Report annotation rendered by the rule-captions report section. Inline text: report = "# QC summary\nReads trimmed with fastp." (markdown allowed). Or a table: report = { file = "notes/qc.md", caption = "Alignment QC" } where file is a markdown/text file resolved relative to the workdir (at least one of file/caption required) — see Rule Captions |
threads |
Integer | No | (Deprecated) CPU threads — use resources.threads instead. Still honored, but oxo-flow lint flags it with a W025 suggestion to move it under [rules.resources] |
memory |
String | No | (Deprecated) Memory allocation — use resources.memory instead. Still honored, but oxo-flow lint flags it with a W025 suggestion to move it under [rules.resources] |
resources |
Table | No | Full resource specification (threads, memory, gpu, disk, time_limit, partition, groups) |
Run-path field validation (#523). Rule names (charset: alphanumeric,
_,-,:),resources.memory/memoryformats (<number><G|M|K|T>), thread counts andretry_delayformats are validated when the workflow is parsed — an invalid value fails the run before any expansion or rendering, so values from a third-party[[include]]can never reach the rendered docker/scheduler/shell command line. Design-level checks (GPU/backend pairing,output_patternconsistency) stay at expansion time where their context-aware diagnostics run. |environment| Table | No | Environment specification | |transform| Table | No | Unified scatter-gather operator (split → map → combine) | |when| String | No | Conditional expression — skip rule whenfalse. May call read-only runtime functions (reads_count,wc_lines,file_size,file_exists,regex_extract) that inspect files in the run workdir; see Data-DependentwhenGates | |envvars| Table | No | Dictionary of environment variables to inject | |params| Table | No | User-defined parameters for shell templates | |pre_exec| String | No | Command to run before the main shell command | |on_success| String | No | Command to run after rule succeeds | |on_failure| String | No | Command to run after rule fails (all retries exhausted) | |retries| Integer | No | Number of retry attempts on failure (default: 0) | |interpreter| String | No | Explicit interpreter for script execution | |checkpoint| Boolean | No | Enable checkpoint re-entry after this rule completes (requirescheckpoint_manifest) | |checkpoint_manifest| String | No | Path (relative to workdir,{config.x}-expanded) of the TOML manifest the checkpoint rule writes at runtime to declare new samples | |scatter| Table | No | Fan-out parallel execution over a variable with optional gather | |expand_inputs| Table | No | Cartesian product expansion of input patterns | |input_groups| Array of tables | No | Group files by a pattern wildcard into one per-group instance (the NextflowgroupTuplepattern) — see Input Groups | |priority| Integer | No | Execution priority (higher = runs first; default: 0) | |target| Boolean | No | Default target:run/dry-runwithout an explicit-t/--modulebuilds the rules markedtarget = trueand their transitive producers; with no marked rule the full DAG runs (an explicit-talways wins) | |required| Boolean | No | Pipeline fails if this rule fails, even without downstream deps | |optional| Boolean | No | Rule is skipped (no error) when its inputs don't exist — literal globs count as missing only when they match nothing | |benchmark| String | No | Per-instance benchmark file: after the rule completes, the engine writes its sampled metrics (wall time, peak RSS, memory limit, CPU seconds, retries) as JSON to this path — engine wildcards in the path resolve per instance, and a failed write never fails the rule | |log| String | No | Log file path for rule execution output. The engine creates its parent directory and expands{config.x}/wildcards in the path, but populating the file is the rule's job — typically via2> {log}in the shell, so it carries content only when the tool writes to stderr. A tool that self-logs to its own location leaves the declared log empty; that never means diagnostics were lost — the captured stdout/stderr tails are stored in the checkpoint'srule_runsregardless, and a succeeded rule whose declared log and both captured tails are all empty earns one Info hint that the tool likely self-logs (issue #832) | |group| String | No | Reserved metadata: cluster array grouping keys on the rule's template name by design; this label is round-tripped for external tooling | |cache_key| String | No | Content-addressed output reuse: cached outputs are restored when the key, inputs, outputs, and rendered command hash identically to a previous run (issue #194 §2.3) | |input_function| String | No | Parsed but not yet called — no dynamic input resolution is performed today | |rule_metadata| Table | No | Domain-specific metadata (assay, organism, etc.) — a round-tripped surface for external tooling; the engine itself does not render it | |env_group| String | No | Reference to a named environment in[env_groups]| |depends_on| Array | No | Explicit rule-level dependencies (by rule name) | |extends| String | No | Inherit settings from a base rule | |retry_delay| String | No | Delay between retries (e.g.,"5s","30s","2m") | |temp_output| Array | No | Scratch artifacts the rule itself overwrites (e.g..tmp.bam). Cleaned when the rule fails (so a stale partial never masquerades as a completed run); they persist on success — marktemporary = truefor success-path deletion | |temporary| Boolean | No | Delete the rule's outputs after a fully successful run once every dependent has completed, recording a tombstone so a future run regenerates them on demand (leaf rules keep their outputs) | |scratch| Boolean | No | Execute in an isolated per-instance scratch directory: inputs render as absolute paths, declared outputs written there move back to the main workdir and are verified, and the scratch is removed on success (preserved with a path note on failure). Use for tools that write fixed filenames or pollute the workdir | |protected_output| Array | No | Outputs the engine must never destroy: failure invalidation skips them (no delete, no move-aside to<name>.oxo-failed) andoxo-flow cleanrefuses them with a(protected — skipped)diagnostic, even with--force. Three pattern forms are honored: exact paths,{wildcard}patterns (results/{sample}.bam), and shell globs (results/*.bam—*does not cross/). Advisory in the remaining sense that the rule's own command can still overwrite them on a rerun | |tags| Array | No | Categorization tags (e.g.,["qc", "alignment"]) |
| ancient | Array | No | Inputs excluded from freshness tracking and input manifests — they never trigger re-execution (reference files; an engine wildcard in a declaration matches the expanded input structurally) |
| localrule | Boolean | No | Parsed but not enforced yet — mixed local+cluster execution needs a scheduler-side local executor and live cluster verification; every schedulable rule can currently be submitted. Do not rely on it (see Resource Tuning) |
| format_hint | Array | No | Parsed and validated but not yet used by the engine (planned for I/O optimization) |
| pipe | Boolean | No | Parsed and validated but not yet used by the engine (planned for FIFO streaming; no streaming is performed today) |
| checksum | String | No | Parsed but not yet enforced — provenance checksums (sha256) are computed for all outputs regardless; no per-rule verification happens today |
| resource_hint | Table | No | Resource estimation hints for dynamic scheduling |
Note: When a rule declares outputs, at least one of shell, script, or transform must be provided. If both shell and script are defined, they execute sequentially: shell first, then script.
Environment specification#
Exactly one backend per rule (uncomment the one you need):
[[rules]]
name = "example"
# Conda
environment = { conda = "envs/tools.yaml" }
# # Pixi
# environment = { pixi = "envs/pixi.toml" }
# # Docker
# environment = { docker = "biocontainers/bwa:0.7.17" }
# # Singularity
# environment = { singularity = "docker://biocontainers/bwa:0.7.17" }
# # Python venv
# environment = { venv = ".venv/" }
# # HPC modules
# environment = { modules = ["gcc/11.2.0", "openmpi/4.1.1"] }
# # Mamba / micromamba (auto-detects binary, same YAML format as conda)
# environment = { mamba = "envs/qc.yaml", mamba_prefix = ".oxo-flow/envs" }
# # Conda with custom prefix
# environment = { conda = "envs/qc.yaml", conda_prefix = ".oxo-flow/envs" }
# # venv with custom requirements file
# environment = { venv = ".venv/", venv_requirements = "envs/dev-requirements.txt" }
# # Reference a named environment group (defined in [env_groups])
# env_group = "qc_env"
Container shell: docker/singularity rules run under bash inside
the container when the image provides it (falling back to sh otherwise).
nf-core-derived images ship bash, and their scripts rely on bash features
(set -o pipefail, [[ ]]) that the container default sh (often dash)
rejects — the re-exec shim keeps these scripts working unmodified.
Inline conda package lists: besides a YAML path, a conda/mamba spec may be a channel-qualified package list — the concise form for rules whose environment is a handful of pinned tools:
Whitespace- or comma-joined packages are the same list. Every package must
be channel-qualified (channel::name=version) and match the strict package
charset — the tokens are rendered as separate conda create arguments, so
shell metacharacters are still rejected (E016). The env name is derived
from the first package plus a content hash of the whole list: identical
package sets share one environment, a changed set builds a fresh one.
Specs without :: (paths, lockfiles) keep the YAML-path semantics.
Singularity spec shapes: a singularity spec is either a pull URI
(docker://, library://, oras://, https://) — the engine pulls and
converts it to a SIF in the workdir on first use, reusing the SIF when it
exists — or a local SIF path (e.g. a site image store's
/data/images/bwa_0.7.17.sif), which is used as-is with no pull step. A
local path that does not exist fails the rule with a clear diagnostic.
GPU passthrough (gpus): environment = { docker = "nvcr.io/nvidia/clara-parabricks:4.4.0-1", gpus = "all" }
passes the host's GPU devices to the container via docker's
--gpus <value> (also "0" or "device=0,1"). Requires
nvidia-container-toolkit on the host. Declaring gpus without a docker
image is a validation error — the flag would otherwise be silently
ignored on system/conda backends. (Singularity --nv passthrough is not
implemented yet; gpus with a singularity image is rejected for now.)
Named Environment Groups ([env_groups])#
Instead of repeating the same environment spec across multiple rules, define
named groups once in [env_groups] and reference them via env_group:
[env_groups.qc_env]
conda = "envs/qc.yaml"
[env_groups.align_env]
conda = "envs/alignment.yaml"
[[rules]]
name = "fastqc"
env_group = "qc_env"
input = ["raw/{sample}.fastq.gz"]
output = ["qc/{sample}_fastqc.html"]
shell = "fastqc {input} -o qc/"
Rules using env_group inherit the full environment specification from the
named group. The env_group spec takes precedence over an inline
[rules.environment]; the inline spec is only used when no env_group is set
(or the named group does not exist), followed by [defaults] environment.
Environment Variables (envvars)#
Inject rule-specific environment variables directly into the execution context:
[[rules]]
name = "deep_learning"
shell = "python train.py"
[rules.envvars]
CUDA_VISIBLE_DEVICES = "0"
PYTHONPATH = "./src"
Variables defined here are available to the main shell command as well as all lifecycle hooks (pre_exec, etc.).
Parameters (params)#
Define custom variables for use in shell templates. Unlike [config], which is global, params are specific to a single rule and take precedence during interpolation:
[[rules]]
name = "count_reads"
shell = "samtools view -c -q {params.min_qual} {input} > {output}"
[rules.params]
min_qual = 20
Script Execution (script)#
The script field allows you to execute external script files (Python, R, etc.) with automatic interpreter detection.
[[rules]]
name = "analyze"
script = "scripts/analysis.py --min-quality {params.q}"
interpreter = "python3" # Optional: overrides auto-detection
Interpreter Detection Order:
- Explicit
interpreterfield on the rule. - Custom
[workflow.interpreter_map]in the metadata. - Built-in defaults based on file extension.
Lifecycle Hooks#
Hooks allow you to run auxiliary logic at different stages of a rule's life:
[[rules]]
name = "process_data"
shell = "python process.py"
pre_exec = "mkdir -p tmp_workspace"
on_success = "echo 'Success!' | slack-notify"
on_failure = "rm -rf tmp_workspace && echo 'Cleanup done'"
retries = 3
pre_exec: Runs before the main command. If it fails, the rule is aborted.on_success: Runs only after the main command completes with exit code 0.on_failure: Runs if the main command fails, after allretrieshave been exhausted.
Hook commands support the same {config.x} / {input} / {output} placeholder expansion as shell.
Output validation: After a rule's shell command exits successfully (exit code 0),
oxo-flow verifies that all declared output files exist on disk. If any declared output
is missing — for example, because a multi-step shell script's cleanup step masked an
earlier tool failure — the rule is treated as failed (with exit code sentinel -1)
and the on_failure hook runs instead of on_success. This prevents silently broken
rules from being checkpointed as "completed" and skipped on resume.
Resources (extended)#
For rules needing GPU, disk, or time limits, use the resources sub-table:
[[rules]]
name = "gpu_task"
input = ["data.h5"]
output = ["model.pt"]
shell = "python train.py"
[rules.resources]
threads = 8
memory = "64G"
gpu = 1
disk = "200G"
time_limit = "48h"
Declared disk requirements are checked against the working directory's
free space before a run starts — a shortfall prints a warning (the run
proceeds; the warning exists so long jobs don't fail mid-pipeline).
| Field | Type | Example | Description |
|---|---|---|---|
threads |
Integer | 8 |
Number of CPU threads |
memory |
String | "16G" |
Memory allocation |
gpu |
Integer | 1 |
Number of GPUs |
disk |
String | "200G" |
Local disk space |
time_limit |
String | "48h" |
Wall-time limit |
partition |
String | "gpu" |
HPC partition/queue to submit to |
groups |
Table | {db_conn = 1} |
Resource group consumption tracking |
Resource Management#
Declaration vs Enforcement#
oxo-flow tracks declared resources for scheduling but does not strictly enforce them in local execution. On HPC clusters, resources are enforced by the scheduler.
Local execution:
- Resources are tracked to prevent over-allocation
- Warnings emitted when declaring resources exceeding system capacity
- Jobs may oversubscribe if user intentionally requests more than available
HPC clusters:
- Resources translated to scheduler directives (SLURM, PBS, SGE, LSF)
- Scheduler enforces limits - jobs requesting more than allocated will fail
Platform Detection#
| Platform | Thread Detection | Memory Detection |
|---|---|---|
| Linux | num_cpus crate |
sysinfo crate |
| macOS | num_cpus crate |
sysinfo crate |
Validation Warnings#
When a rule declares resources exceeding system capacity, oxo-flow emits warnings during validation but does not block execution:
⚠️ rule 'bwa_align' requests 128 threads but system has 64 (will oversubscribe)
⚠️ rule 'big_sort' requests 128GB but system has 32GB (may OOM)
This allows intentional oversubscription for testing or when user knows better.
Cleanup Behavior#
oxo-flow automatically cleans up temporary outputs:
| Scenario | Cleanup |
|---|---|---|
| Success + temp_output | Kept — no success-path cleanup runs; mark temporary = true instead if you want outputs deleted after the run |
| Failure + temp_output | Cleaned to prevent stale partial files |
| Failure + declared output | Partial outputs created by the failed attempt are deleted; all pre-existing declared outputs are moved aside as <name>.oxo-failed (modified or untouched — a stale file from an earlier run era would otherwise block the re-execution the recorded failure schedules with File exists errors; issue #756) |
| Transform with cleanup=true | Chunk files cleaned after the whole run finishes successfully (kept on failed runs for debugging; re-runs recompute the map rules) |
| Success + temporary = true | Outputs deleted after the run once every dependent rule has completed; a tombstone is recorded in the checkpoint so a later run that needs the outputs regenerates the rule first (lazy cascade-up). Leaf rules (no dependents) keep their outputs. For output_pattern rules, deletion is per recorded instance (only the concrete paths that ran and completed — issue #844); patterns that cannot be expanded produce a warning naming the rule instead of a silent no-op. If a dependent has not completed yet, the outputs are kept and a warning names the rule and its pending dependents — a later run (or resume) cleans them up once the dependents finish |
temporary is for whole intermediates (e.g. multi-GB per-sample BAMs kept
only until the queue-level callers finish): the deletion is checkpoint-aware,
so a plain re-run skips the rule and does NOT regenerate the file — it comes
back only when a dependent actually needs it again.
Failed-rule output invalidation. A rule that fails mid-write must never
leave partial outputs that the freshness gate would mistake for a completed
run. Before executing, the engine snapshots each declared output's existence
and modification time; when the rule fails (non-zero exit, timeout, or
missing outputs after an exit-0 shell), the declared outputs are
invalidated: files created during the attempt are deleted, and every
pre-existing declared output moves aside as <name>.oxo-failed — modified
or untouched (issue #756). The recorded failure schedules a re-execution,
and a stale file surviving at the declared path would block that re-run
inside the script (ln: File exists, mv: cannot move ... File exists);
move-aside keeps it recoverable, and protected_output remains the escape
hatch for files that must survive any failure. The next run then
re-executes the rule instead of skipping it as up-to-date. Note this
covers failures the engine observes; a hard kill (SIGKILL) mid-rule can
still leave outputs behind, since no cleanup code runs — the --rerun
flag is the escape hatch for that case.
Timeout Enforcement#
On Unix systems (Linux, macOS), a timeout kill walks the rule's process subtree rather than a process group — rules deliberately run inside the run's process group ("one run = one process group"), so the engine snapshots parent→child links via sysinfo and signals each descendant individually, deepest first: SIGTERM to the whole subtree, a 10-second grace poll for survivors, a re-scan for descendants spawned during the window, then SIGKILL to whatever remains (issue #194):
GPU Specification#
For detailed GPU requirements:
[rules.resources.gpu_spec]
count = 2
model = "A100" # SLURM: --gres=gpu:a100:2
memory_gb = 40 # SLURM: --mem-per-gpu=40G
compute_capability = "8.0" # For filtering (not scheduler directive)
Note: PBS/SGE GPU syntax varies by site. Use extra_args for site-specific flags.
Resource Hints#
When exact requirements unknown, provide hints for estimation:
[rules.resource_hint]
input_size = "medium" # small (~1GB), medium (~10GB), large (~100GB), xlarge (~500GB)
memory_scale = 2.0 # Estimated memory = input_size × scale
runtime = "slow" # reserved: parsed, not yet consumed by the scheduler
io_bound = true # reserved: parsed, not yet consumed by the scheduler
Memory estimation formula: estimated_mb = input_size_mb × memory_scale. The
runtime and io_bound hints are reserved scheduler knobs: only
input_size and memory_scale are consumed today.
Script Execution#
shell commands run through bash (falling back to sh when bash is
unavailable), so bashisms — brace expansion, [[ ]] conditionals, process
substitution, $(( )) arithmetic — work in rule shells.
Script Field#
Execute a script file instead of (or in addition to) a shell command:
[[rules]]
name = "analysis"
input = ["data.csv"]
output = ["results.json"]
script = "scripts/analyze.py" # Auto-detects interpreter from extension
When both shell and script are defined, they execute sequentially: shell first, then script.
[[rules]]
name = "qc_and_report"
shell = "fastqc {input} -o qc/"
script = "reports/qc_report.qmd" # Runs after shell completes
Interpreter Detection#
oxo-flow automatically detects the interpreter from script file extension:
| Extension | Interpreter | Notes |
|---|---|---|
.py |
python |
Python script |
.R / .r |
Rscript |
R script |
.jl |
julia |
Julia script |
.sh / .bash |
bash |
Shell script |
.pl |
perl |
Perl script |
.rb |
ruby |
Ruby script |
.qmd |
quarto render |
Quarto document |
.Rmd / .rmd |
quarto render |
R Markdown |
.ipynb |
jupyter nbconvert --to notebook --execute |
Jupyter notebook |
.smk |
snakemake |
Snakemake workflow |
.nextflow |
nextflow run |
Nextflow script |
.wdl |
miniwdl run |
WDL workflow |
Explicit Interpreter Override#
Override auto-detection with interpreter field:
[[rules]]
name = "custom_python"
script = "analyze.py3"
interpreter = "python3.11" # Override default python
Custom Interpreter Map#
Configure custom interpreter mappings at workflow level:
[workflow]
name = "pipeline"
[workflow.interpreter_map]
".m" = "octave" # MATLAB/Octave
".sas" = "sas" # SAS
".do" = "stata-mp" # Stata
".stan" = "cmdstan" # Stan
Additional Rule Fields#
Output Management#
| Field | Type | Description |
|---|---|---|
temp_output |
Array | Scratch artifacts cleaned when the rule fails; persist on success (see Cleanup Behavior) |
protected_output |
Array | Outputs the engine never deletes or moves aside on failure; clean also refuses them. Exact paths, {wildcard} patterns, and shell globs are all honored (see the rule-field table above) |
temporary |
Boolean | Delete the rule's outputs after a fully successful run once every dependent has completed (tombstone + lazy regeneration; leaf rules keep outputs) |
[[rules]]
name = "align"
output = ["aligned/{sample}.bam", "aligned/{sample}.bam.bai"]
temp_output = ["aligned/{sample}.tmp.bam"] # Cleaned if the rule fails; kept on success
temporary = true # Delete aligned/*.bam once all callers finish
Directory outputs#
Directories are valid outputs — declare them with a trailing slash
(or a bare directory name). Everything a rule writes inside that
directory counts as produced, which is the idiomatic way to capture
tool-generated side directories such as multiqc_data/ and
multiqc_plots/:
[[rules]]
name = "multiqc"
output = [
"results/multiqc/multiqc_report.html",
"results/multiqc/multiqc_data/",
"results/multiqc/multiqc_plots/",
]
Declaring directories matters for two reasons: downstream rules and the report see the whole directory as available, and the executor checks declared outputs after the shell exits — a rule that exits 0 while a declared output (file or directory) is missing is marked as failed. Conversely, do not declare directories the shell only uses as scratch: an undeclared directory that the rule recreates on re-run is lost on resume.
Execution Control#
| Field | Type | Default | Description |
|---|---|---|---|
depends_on |
Array | — | Explicit rule dependencies (not inferred from files) |
localrule |
Boolean | false |
Parsed but not enforced yet — every schedulable rule can currently be submitted to the cluster |
| checkpoint | Boolean | false | Enable checkpoint re-entry (requires checkpoint_manifest; the rule must not use {sample}/{group}) |
DAG edge inference — edges form when a rule's input paths match
another rule's output paths, for every input form: template paths,
expand_inputs patterns (both the expanded concrete paths at run time
and the raw patterns in the template-level graph — multiqc-style
aggregators therefore connect to their contributors even in
graph -f dot without --expanded), glob patterns
(raw/*.fastq.gz), and directory inputs (every producer writing under
that directory). Rules that declare the SAME output string (shared
bins/staging directories — e.g. two assembler variants emitting into one
GenomeBinning/<tool>/bins dir) link ALL of their producers: every
consumer of the shared path is ordered after every writer.
depends_on remains available for ordering that files cannot express,
and duplicate edges are deduplicated.
[[rules]]
name = "setup"
shell = "mkdir -p results"
depends_on = [] # Run first, before file-based dependencies
[[rules]]
name = "local_only"
shell = "echo 'local task'"
localrule = true # NOTE: not yet enforced — every schedulable rule can be submitted
Input Groups (input_groups)#
Group many files of one sample (or any pattern key) into a single
per-key rule instance whose {input} renders all of them — the Nextflow
groupTuple pattern. Unblocks lane-merging (cat a sample's lane files
in one step), library merging, and multi-file per-sample catalog steps.
[[rules]]
name = "lanemerge"
input_groups = [
{ pattern = "results/adapterremoval/{sample}_{lane}_R1.fastq.gz",
group_by = "sample",
keep = "lane" }
]
output = ["results/merged/{sample}_R1.fastq.gz"]
shell = "cat {input} > {output}"
pattern— path glob with{wildcard}placeholders, relative to the workflow file's directory (the same rootsample_patternscans). A bare*outside{wildcard}spans is a segment-local glob (issue #246):raw/{sample}_*_R1.fastq.gzgroups every lane file of a sample without enumerating lane names — the CAT_FASTQ shape, where the per-sample file count is not known ahead of time and no[[values]]table can list it.**(cross-segment recursion) is rejected. Note a{wildcard}span may span a separator by design (\S+); the segment locality applies to bare*only.group_by— the wildcard that names each group. One instance is created per discovered group value. May also be a metadata column,group_by = "meta.<column>"— see below.keep— optional restriction: which other pattern wildcards are bound to the group's first (sorted) file and become available as{wildcard}placeholders inoutput/shell. Omitted = all other wildcards. A string or an array of strings is accepted. Not allowed withgroup_by = "meta.<column>"(metadata-grouped instances bind only the column name).{input}— all matched files of the group, space-joined in stable sorted order (files sort by path; the first sorted file supplies thekeepbindings, so results never depend on readdir order).{input_group.<wildcard>}— a captured per-group value as a list (space-joined), e.g.{input_group.lane}rendersL1 L2. These are lists: use them in shell loops/commands, not as single paths.- Files are discovered at plan time from two sources: files already
on disk under the workflow root, and the literal outputs of upstream
rules — so a producer rule declared earlier in the file (e.g. a
per-
{sample}_{lane}rule) flows straight into the group's{input}, with DAG edges inferred from those plan-time-known paths. - Regular
inputentries coexist: group files render first, declared inputs append after them in{input}. - A group key with zero files instantiates nothing — the rule is skipped for that key with a warning (no instance = nothing to run).
- When the workflow declares a sample domain (
[[sample_groups]],sample_pattern, or[[pairs]]),group_by = "sample"group keys must be in that domain — keys outside it are pruned with a warning (#246). A key whose value is an underscore-suffixed form of a declared sample (S1_REP1for declaredS1) is kept: replicate-suffix naming folds by base id, the same convention the group pattern itself uses.
v1 constraints: at most one input_groups entry per rule;
** recursion is not supported in pattern (bare * globs within one
segment are supported);
producers must be declared before their input_groups consumer;
the keep wildcard bindings come from the first sorted file only (the
captured per-group values of every file stay available via
{input_group.*}); {config.x} placeholders in pattern resolve
against the workflow config. A producer whose output is a directory
pattern (e.g. results/star/{sample}/) contributes only the directory
path itself to the candidate pool — files inside it do not match
input_groups patterns (declare the producer's files as outputs to
group them).
# A multiqc-style aggregator over grouped files: one instance per
# sample, receiving every fastQC zip of that sample.
[[rules]]
name = "multiqc"
input_groups = [
{ pattern = "qc/fastqc_{sample}_{replicate}.zip", group_by = "sample" }
]
output = ["qc/multiqc_{sample}.html"]
shell = "multiqc {input} -o qc -n {output}"
Grouping by metadata column#
group_by = "meta.<column>" groups the matched files by a
metadata_file column value instead of
a pattern wildcard — the chipseq multi-antibody pattern, where the BAMs of
several samples carry the same antibody and one peak-calling instance per
antibody is wanted:
[workflow]
metadata_file = "samples.tsv" # sample antibody
# S1 H3K27ac
# S2 H3K27ac
# S3 Input
[[rules]]
name = "peaks"
input_groups = [
{ pattern = "results/mapping/{sample}.bam", group_by = "meta.antibody" }
]
output = ["peaks/{antibody}.bed"]
shell = "macs2 callpeak -t {input} -n {antibody} --outdir peaks && echo {input_group.sample}"
- One instance per distinct column value:
peaks_H3K27acreceivesresults/mapping/S1.bam results/mapping/S2.bamandpeaks_Inputreceivesresults/mapping/S3.bam. - The group key binds under the column name (
{antibody}), not{sample}— a metadata-grouped instance has no single sample. Files whose row has an empty column value are skipped. - Outputs must not reference pattern wildcards (
{sample}in an output is a plan-time error); the per-group sample names stay available as{input_group.sample}. keepis rejected withgroup_by = "meta.<column>"— the column name is the instance's only binding.
Retry Configuration#
| Field | Type | Default | Description |
|---|---|---|---|
retries |
Integer | 0 | Number of automatic retry attempts |
retry_delay |
String | — | Delay between retries ("5s", "30s", "2m") |
[[rules]]
name = "network_task"
shell = "curl https://api.example.com/data"
retries = 3
retry_delay = "30s"
Input/Output Hints#
| Field | Type | Description |
|---|---|---|
ancient |
Array | Inputs excluded from freshness tracking and input manifests — they never trigger re-execution |
format_hint |
Array | Parsed but not yet used by the engine (planned for I/O optimization) |
pipe |
Boolean | Parsed but not yet used by the engine (planned for FIFO streaming; no streaming is performed today) |
checksum |
String | Parsed but not yet enforced — provenance checksums (sha256) are computed for all outputs regardless |
[[rules]]
name = "align"
input = ["reads/{sample}.fastq.gz", "ref/hg38.fa"]
ancient = ["ref/hg38.fa"] # never triggers rebuild, however new the file is
# format_hint and checksum are accepted but currently have no effect
Organization#
| Field | Type | Description |
|---|---|---|
tags |
Array | Categorization tags (["qc", "alignment"]) |
extends |
String | Base rule to inherit settings from |
[[rules]]
name = "align_default"
tags = ["alignment", "production"]
[rules.resources]
threads = 8
memory = "32G"
[[rules]]
name = "align_fast"
extends = "align_default" # Inherits threads, memory, tags
[rules.resources]
threads = 16 # Override inherited value
Priority and Targeting#
| Field | Type | Description |
|---|---|---|
priority |
Integer | Execution priority (higher runs first; default: 0) |
target |
Boolean | Default target: built by run/dry-run when no explicit -t/--module is given (marked rules + their transitive producers; none marked → full DAG) |
[[rules]]
name = "critical_step"
priority = 10 # Runs ahead of lower-priority rules
target = true # built by `run`/`dry-run` when no -t is given
Optional and Required Rules#
| Field | Type | Description |
|---|---|---|
optional |
Boolean or "any" |
true: the rule is skipped (no error) when ANY declared input is missing — literal globs count as missing only when they match nothing. "any": the rule runs when AT LEAST ONE declared input exists and is skipped only when none do — the alternative-input pattern, e.g. a consumer that accepts whichever mapper/aligner wrote the file. false (default): every input must exist |
required |
Boolean | If false, a failure of this rule does not fail the run: the failure is recorded and listed, its dependents are blocked, and the run exits 0 (best-effort / continue-on-error, e.g. QC steps). Defaults to true — required failures fail the run (or, with --keep-going, keep it running but still exit non-zero with the failures listed: keep-going changes scheduling, never the verdict) |
[[rules]]
name = "experimental"
optional = true # Skip if input data is absent
required = false # Failure is reported but does not fail the run
[[rules]]
name = "filter_bam"
# Run when any of the mapper-specific BAMs exists (bwa writes *_PE.mapped.bam,
# bwamem/bt2/circularmapper write {sample}.mapped.bam); skip only when none do.
optional = "any"
input = [
"results/mapping/bwa/{sample}_PE.mapped.bam",
"results/mapping/bwa/{sample}.mapped.bam",
"results/mapping/bt2/{sample}.mapped.bam",
]
When a rule is skipped because none of its optional inputs exist, non-optional dependents are skipped too ("blocked by skipped upstream dependency") instead of executing a doomed shell command on missing files.
Logging and Benchmarking#
| Field | Type | Description |
|---|---|---|
log |
String | File path for capturing rule stdout/stderr |
benchmark |
String | Per-instance benchmark file — the engine writes its sampled metrics as JSON here after the rule completes |
[[rules]]
name = "align"
log = "logs/align_{sample}.log"
benchmark = "benchmarks/align_{sample}.json" # per-instance JSON, written by the engine
Job Grouping and Caching#
| Field | Type | Description |
|---|---|---|
group |
String | Reserved metadata: cluster array grouping keys on the rule's template name by design; this label is round-tripped for external tooling |
cache_key |
String | Content-addressed output reuse: cached outputs are restored when the key, inputs, outputs, and rendered command hash identically to a previous run (issue #194 §2.3) |
[[rules]]
name = "variant_call"
group = "variant_calling" # reserved metadata (grouping keys on the template name)
cache_key = "vc_v2.0" # Content cache key — outputs are reused when content is identical
Dynamic Input Resolution#
| Field | Type | Description |
|---|---|---|
input_function |
String | Parsed and included in the rule fingerprint but not yet called — no dynamic input resolution is performed today (use expand_inputs or output_pattern for dynamic inputs) |
Arbitrary Metadata#
| Field | Type | Description |
|---|---|---|
rule_metadata |
Table | Domain-specific metadata (assay type, organism, protocol, etc.) — a round-tripped surface for external tooling; the engine itself does not render it |
[[rules]]
name = "wgs_align"
[rules.rule_metadata]
assay = "WGS"
organism = "Homo sapiens"
protocol = "Illumina_NovaSeq_6000"
Scatter-Gather (Legacy)#
The scatter field provides fan-out parallelism over a variable with optional
gather. For new workflows, prefer the unified
transform operator.
| Field | Type | Description |
|---|---|---|
scatter.variable |
String | Variable to scatter over (e.g., "chr") |
scatter.values |
Array | Values to scatter across |
scatter.values_from |
String | Config variable reference for values |
scatter.gather |
String | Name of the gather rule |
A rule whose scatter fans out over zero values — declared values = [], or
values_from referencing a config key that resolves to an empty list —
produces zero instances and can never run. oxo-flow validate therefore
skips the input-existence check (E010/W020) for such rules and excludes their
output templates from the W033 output-collision lint, mirroring the
when-gated-off exemption: a rule that cannot run must not fail validation
over inputs it will never read. An unresolvable values_from (missing
config key, non-string/array value) is deliberately NOT treated as empty —
the executor expands zero instances silently in that case, so validate keeps
the input-existence check to surface the typo as a missing-input diagnostic
instead of silently passing (issue #616; the stricter [[values]] tables and
transform.split reject an unresolvable reference outright).
Expand Inputs#
The expand_inputs field generates additional input combinations via Cartesian
product expansion.
| Field | Type | Description |
|---|---|---|
expand_inputs[].pattern |
String | Input pattern with variables |
expand_inputs[].variables |
Table | Variable name → config reference (TOML array or comma-separated string) |
[[rules]]
name = "multi_ref_align"
expand_inputs = [
{ pattern = "refs/{ref_genome}.fa", variables = { ref_genome = "config.ref_genomes" } }
]
Variables may also reference the [config] section — either a TOML array
or a string:
[[sample_groups]]
name = "cohort"
samples = ["NA12878", "NA12879", "NA12880"]
[[rules]]
name = "combine_gvcfs"
input = []
expand_inputs = [
{ pattern = "variants/{sample}.g.vcf.gz", variables = { sample = "config.samples_list" } }
]
Config reference semantics:
| Config value | Expands to |
|---|---|
["a", "b"] (array) |
a, b |
"a" (string, no comma) |
a |
"a,b,c" (comma-joined string) |
a, b, c (split on commas, trimmed) |
["a,b"] (single-element array) |
a,b as one value — escape hatch for comma-containing strings |
Inline literals (a plain string written directly in variables, e.g.
sample = "S1,S2") are always treated as one value and are never split —
the comma-splitting applies only to config.* references.
Comma-joined strings are how the engine injects merged sample lists
(config.samples_list, config.samples_<group_name> — see
Merging Multiple Sample Sources), so
they can be referenced directly here.
The aggregation idiom — multi-sample merge/collapse steps consume the
expanded inputs through {input}, which renders space-joined:
[[rules]]
name = "merge_bams"
expand_inputs = [
{ pattern = "aln/{sample}.bam", variables = { sample = "config.samples_list" } }
]
output = ["merged.bam"]
shell = "samtools merge {output[0]} {input}" # aln/S1.bam aln/S2.bam ...
Shell loops over the same list work identically
(for f in {input}; do ...; done — the pattern used by the multiqc
aggregation rules across the community catalog). Note the deliberate
split: config.samples_list is comma-joined for expansion-time
fan-out, while {input} is space-joined for shell consumption. Do
not shell-loop over {config.samples_list} directly — a comma-joined
value iterates as a single word; use tr ',' ' ' only if the expansion
form is genuinely unavailable. Chunk-level aggregation is the transform
operator's combine stage instead (see
Transform Operator).
Namespace collision with [[values]] tables — a bare {name} in an
expand_inputs pattern that matches a [[values]] table name is treated as
that table: the rule fans out once per table entry, and the pattern is
expanded entry-by-entry ({chr} with values = ["chr1", …, "chrM"] means
24 separate instances, not "all chromosomes at once"). Writing an
aggregation rule that reuses a table name for its literal meaning
("collect every chr's calls") therefore silently produces a per-entry
scatter whose instances usually share one output path and race — the
repair is a distinct variable name for the aggregation reading, or keying
the outputs by the table's wildcard for a genuine per-entry rule:
[[values]]
name = "chr"
values = ["chr1", "chr2", "chr3"]
[config]
chrom_list = ["chr1", "chr2", "chr3"] # the aggregation binding's source
# WRONG: 'chr' is the [[values]] table — the rule fans into 3 instances
# per pair, every one writing the SAME vcf.raw/{pair_id}.vcf.gz and racing
[[rules]]
name = "merge_chr_vcfs_bad"
expand_inputs = [{ pattern = "vcf.call/{chr}/{pair_id}.vcf.gz" }]
output = ["vcf.raw/{pair_id}.vcf.gz"]
# RIGHT (aggregation): bind the dimension through `variables` instead —
# the pattern expands to ALL entries inside ONE instance ({input} is the
# space-joined list), so the merge happens once
[[rules]]
name = "merge_chr_vcfs"
expand_inputs = [
{ pattern = "vcf.call/{chrom}/{pair_id}.vcf.gz",
variables = { chrom = "config.chrom_list" } }
]
output = ["vcf.raw/{pair_id}.vcf.gz"]
# RIGHT (scatter): keep the bare table wildcard, but key the outputs by it
[[rules]]
name = "call_chr"
output = ["vcf.call/{chr}/{pair_id}.vcf.gz"]
The two "RIGHT" forms differ in cardinality: the variables-bound
aggregation yields one instance whose {input} lists every chromosome;
the bare-wildcard scatter yields one instance per chromosome (and
typically a downstream aggregation rule over the keyed outputs).
The preflight reports the racing shape as SCI-AGG-RACE, naming the
unkeyed fan-out dimension (see
oxo-flow dry-run → Scientific Preflight).
Wildcards#
Wildcards enable dynamic, pattern-based pipeline definitions. For a detailed guide on how they are discovered, expanded, and constrained, see the Wildcards Reference.
Basic Syntax#
Use {name} in file paths for dynamic expansion:
Built-in Placeholders#
Built-in placeholders use the same syntax but have reserved meanings:
| Placeholder | Expands to |
|---|---|
{input} |
Space-separated list of all input files |
{input[N]} |
The Nth input file (0-indexed) |
{input.name} |
The value of the name entry in the rule's input map (input = { name = "..." } → {input.name}); any key works: input = { reads = "..." } → {input.reads} |
{output} |
Space-separated list of all output files |
{output[N]} |
The Nth output file (0-indexed) |
{output.name} |
The value of the name entry in the rule's output map (output = { name = "..." } → {output.name}); any key works: output = { bam = "..." } → {output.bam} |
{threads} |
Thread count assigned to this rule |
{memory} |
Memory allocation assigned to this rule |
{effective_threads} |
Declared threads clamped to the machine's CPUs — the tool-facing concurrency. Use it for flags like --threads/-t when the declared value is an HPC-scale label (e.g. a rule declaring threads = 12 renders 4 on a 4-core box). Unset threads render as 1 |
{effective_memory_mb} |
Declared memory clamped to the machine's total memory, as an integer number of MB — the tool-facing heap budget. Use it for JVM flags (-Xmx{effective_memory_mb}m, -Xms…) so tools size themselves to the box instead of OOM-ing on hardcoded HPC values (-Xmx29491M class). Unset memory renders as the machine's full total |
{log} |
Path of the rule's log field (every instance wildcard — {sample}, {pair_id}, {assembler} … — and {config.x} are expanded per instance; the parent directory is created automatically) |
{config.*} |
Value from the [config] section (plain value, declared default, or CLI override) |
{wildcard.*} |
Per-instance binding under the same grammar as when conditions — sample-group metadata keys (e.g. {wildcard.umi1_pattern}) and expansion keys. Unbound keys render as their literal brace token |
Why the effective pair exists — {threads}/{memory} render the
declared values, which keep the pool semantics (an over-capacity rule
reserves the whole pool and runs alone). The {effective_*} pair renders
what the tool can actually get on this machine. Rules that embed
resource-sized flags should use the effective pair; rules that only pass
the pool declaration through can keep the plain pair.
Container memory follows the same policy — docker/singularity rules
get the declared memory as their --memory cgroup limit, clamped to the
machine's total the same way (a rule declaring memory = "72G" on a 4G
box runs with --memory 4G, not a limit the kernel can never back).
cgroup-aware tools (picard, STAR, Cell Ranger) therefore see an honest
limit on every machine size.
The ceiling counts swap — the machine total above is RAM + swap
(swap is backable memory the kernel uses under pressure; ignoring it
left real capacity unused on small boxes). Pass --max-memory to pin
the ceiling to RAM only when latency matters more than headroom.
Existing environments are verified before setup — when the
environment cache is cold (e.g. after --rerun wipes the checkpoint)
but the conda environment already exists, oxo-flow verifies it in place
and marks it ready instead of re-running the setup, whose conda env
update fallback re-resolves every dependency over the network.
Environment creation is serialized per environment, not per machine
— concurrent processes creating the same env would corrupt each
other's transaction, so setup holds an exclusive lock keyed by the env
cache key (~/.oxo-flow/locks/env-<sha256>.lock); different envs are
independent by conda/pixi semantics and never contend. The wait is
bounded (2h by default, OXO_ENV_LOCK_TIMEOUT_SECS to tune): the
waiter logs every 60s and a timeout fails the rule with the lock path
in the diagnostic instead of hanging the run. The holder's pid is
written into the lock file for diagnosis. On shared accounts, one
user's slow solve only blocks creators of that same environment.
Spawn watchdog — a rule can in principle sit in the Running state
without ever spawning a child process (a lost scheduler wakeup, an
environment still being set up, or waiting on the resource pool).
Every 30s the engine compares rules it considers running against the
processes it can actually see; a rule that has been Running for over
OXO_FLOW_SPAWN_WATCHDOG_SECS (default 600) with no child process at
all is reported by name, and the report distinguishes two cases
(issue #847). If the rule is queued in the resource pool — it has not
acquired its threads/memory yet — the report is a one-shot WARN
pointing at the scheduler's own 60-second blockage lines, because a
queued wait is benign and self-resolving. A rule that is not queued
(has acquired resources, or is parked in environment setup) and still
has no child process keeps the tracing::error! report — that is the
genuine lost-wakeup signature (issue #685). Both reports land in the
run log as well as stderr. The watchdog only logs — it never kills;
if a named rule never starts, capture the run log and file an issue.
Relatedly, the in-process environment-setup mutex now
waits on the same bounded schedule as the cross-process lock
(OXO_ENV_LOCK_TIMEOUT_SECS, logging every 60s) so a task stuck behind
another task's env creation becomes visible instead of hanging silently.
Array-valued {config.*} — a [config] key holding an array renders
as a space-joined list in the shell (["a", "b"] → a b), matching the
{input} convention; iterate with for x in {config.tools}.
{input} vs {input[N]} — both forms are equivalent when the array has a single entry ({input} joins all inputs with spaces; {input[0]} takes the first). Practical guidance:
- Single input/output: either form works —
{input}and{output}are the simplest - Multiple inputs/outputs: use
{input[0]},{input[1]}… to select specific files, or{input}to pass all of them at once - The indexed form communicates intent ("this rule expects exactly this file"), so it remains useful even for single-element arrays in multi-step pipelines
An index at or beyond the declared count is never substituted: rendering leaves the literal {input[N]} text in the command and the tool fails with an unrelated "No such file" error. oxo-flow lint catches this statically as W034 (which also flags map-form rules whose sorted-key indexing has shifted under a declaration edit — the index positions are the sorted keys, not the text order).
Named Input & Output#
For complex rules with many files, give the input/output fields an
inline map (FilePatterns::Map) and reference the entries with
{input.<name>} / {output.<name>}:
[[rules]]
name = "align"
input = { reads1 = "raw/{sample}_R1.fastq.gz", reads2 = "raw/{sample}_R2.fastq.gz" }
output = { bam = "aligned/{sample}.bam" }
shell = "bwa mem {input.reads1} {input.reads2} > {output.bam}"
There is no [rules.named_input] / [rules.named_output] TOML table —
the map is the value of input/output itself (a sub-table named
named_input fails validation with E017).
A named reference to a key the rule does not declare is likewise never
substituted (literal brace text at render time) and is flagged by W034
along with the rule's actual keys, so a typo like {input.tumer} is
reported at lint time instead of surfacing as a tool failure.
Custom Wildcards#
Any {name} pattern not matching a built-in placeholder is treated as a wildcard. oxo-flow expands these based on:
- File discovery: Scanning for matching files in the
inputpath. - Explicit lists: Defined in
[[pairs]]or[[sample_groups]]. - Parameter lists: Defined in
[[values]].
[[values]] — Parameter Fan-out (WC-03)#
[[values]] declares parameter lists that fan rules out the same way
sample groups do — the oxo-flow equivalent of an upstream
"for each assembler × binner" construct, without hardcoding one rule
per combination:
[[values]]
name = "assembler"
values = ["spades", "megahit"]
[[rules]]
name = "assemble"
input = ["reads/{sample}.fq"]
output = ["assemblies/{assembler}/{sample}/contigs.fa"]
shell = "{assembler} -o {output} {input}"
- Every rule referencing
{assembler}(bare form) or{values.assembler}(namespaced form) expands into one instance per value. The reference may live in any trigger field — inputs, outputs, shell,when,expand_inputspatterns, or the rule'soutput_pattern(issue #296): a pattern-only reference fans the producer out per value with the pattern baked per instance, exactly as it already does for{sample}/pair wildcards.
[[rules]]
name = "assemble"
output_pattern = "results/{assembler}/chunks.txt" # fans out per value
shell = "scripts/build_all.sh"
values_from = "config.x" resolves the value list from a config key at
expansion time — a comma string ("u1,u2") or array, the same semantics
as expand_inputs config references. CLI --arg x=u1,u2 or a profile
then drives the fan-out without editing the workflow; takes precedence
over a static values when both are set.
- Multiple [[values]] tables form a cartesian product with each other
and with sample groups / pairs (deterministic order: the first table
varies slowest).
- Instance names follow the existing convention:
assemble_assembler_spades_batch_S1.
- Value groups must have unique names and must not collide with the
reserved wildcards (sample, pair_id, experiment, control,
group, tumor, normal, experiment_type, tumor_type) or the
executor placeholders (input, output, log, threads, memory) —
a table named input would replace the placeholder in every rule's
shell.
[[pairs]] — Experiment-Control Pairing (WC-01)#
[[pairs]] defines experiment-control sample pairs for somatic variant calling and other comparative analyses.
[[pairs]]
pair_id = "CASE_001"
experiment = "EXP_01"
control = "CTRL_01"
[[pairs]]
pair_id = "CASE_002"
experiment = "EXP_02"
control = "CTRL_02"
| Field | Type | Required | Description |
|---|---|---|---|
pair_id |
String | Yes | Unique identifier for this pair |
experiment |
String | Yes | Experiment sample name (alias: tumor) |
control |
String | No | Matched control sample name (alias: normal) |
experiment_type |
String | No | Optional cohort label (alias: tumor_type) |
metadata |
Table | No | Arbitrary key-value pairs (each key becomes a wildcard) |
when |
String | No | Pair-level gate evaluated at plan time: while it evaluates false, the pair declares no rule instances. Uses the same condition vocabulary as rule when (config.<key> truthiness/comparisons, &&/\|\|/!, parentheses, file_exists(...)). A config.<key> reference with no matching [config] key evaluates false (the unbound→false stance) — a plan-time warning flags the referenced key so a typo does not silently drop the pair's whole rule set |
Any rule that references {experiment}, {control}, or {pair_id} in its input, output, or shell fields is automatically expanded into one concrete rule instance per pair. Rules that do not reference any pair wildcard are kept as-is.
Expanded rule naming: {rule_name}_{pair_id} (e.g., mutect2_CASE_001).
Pair-level gating with when#
One static [[pairs]] table can carry profile-switched sample sets — a
toggle in [config] decides which pairs fan out:
[config]
extended = false # flip via --profile or CLI override
[[pairs]]
pair_id = "EXTRA_001"
experiment = "EXP_02"
control = "CTRL_02"
when = "config.extended"
While extended is false, EXTRA_001 fans out no rules; flipping it and
re-running restores the pair's instances (re-expansion from the preserved
templates, the checkpoint re-entry pattern). Unknown keys evaluate false,
so a misspelled key (when = "config.extened") gates the pair out with a
plan-time warning naming the offending key.
Loading pairs from external file#
For large cohort studies with hundreds or thousands of pairs, use pairs_file in [workflow]:
TSV format (tab-separated, header required):
pair_id experiment control experiment_type
CASE_001 EXP_01 CTRL_01 lung_adenocarcinoma
CASE_002 EXP_02 CTRL_02 colorectal
CASE_003 EXP_03 CTRL_03 breast_cancer
CSV format (comma-separated):
pair_id,experiment,control,experiment_type
CASE_001,EXP_01,CTRL_01,lung_adenocarcinoma
CASE_002,EXP_02,CTRL_02,colorectal
JSON format:
[
{"pair_id": "CASE_001", "experiment": "EXP_01", "control": "CTRL_01"},
{"pair_id": "CASE_002", "experiment": "EXP_02", "control": "CTRL_02"}
]
Inline [[pairs]] and pairs_file can be used together; entries from both sources are merged. Each pair_id must still come from exactly one source: a pair_id declared in both with identical content is deduplicated with a warning, while the same pair_id with different experiment/control content is a config error naming both declarations (both copies would fan out rules under the same {rule}_{pair_id} name). The same rule applies to duplicate sample-group names across [[sample_groups]] and sample_groups_file.
Auto-discovery from file pattern#
For workflows with existing paired files, use pairs_pattern in [workflow] to auto-discover pairs by scanning the filesystem:
[workflow]
name = "somatic-calling"
pairs_pattern = "aligned/{pair_id}/{experiment}_vs_{control}.bam"
oxo-flow scans files matching this pattern and extracts wildcards from paths. For a file:
Creates pair:
pair_id = CASE_001experiment = EXP_01control = CTRL_01
Pattern requirements:
- Must contain
{pair_id},{experiment}, and{control}wildcards - Optional
{experiment_type}wildcard also extracted - Pattern is converted to glob (
*) for filesystem scan
This eliminates the need for manual pair lists or external files when working with pre-organized directory structures.
Example#
[[pairs]]
pair_id = "CASE_001"
experiment = "EXP_01"
control = "CTRL_01"
[[rules]]
name = "mutect2"
input = ["aligned/{experiment}.bam", "aligned/{control}.bam"]
output = ["variants/{pair_id}.vcf.gz"]
shell = "gatk Mutect2 -I {input[0]} -I {input[1]} -normal {control} -O {output[0]}"
Produces rule mutect2_CASE_001 with concrete file paths.
See examples/gallery/15_paired_experiment_control_pairs.oxoflow for a full clinical somatic calling pipeline.
[[sample_groups]] — Multi-Sample Cohorts (WC-02)#
[[sample_groups]] organises samples into named groups (e.g., case vs. control) for cohort studies.
[[sample_groups]]
name = "control"
samples = ["CTRL_001", "CTRL_002", "CTRL_003"]
[[sample_groups]]
name = "case"
samples = ["CASE_001", "CASE_002"]
| Field | Type | Required | Description |
|---|---|---|---|
name |
String | Yes | Group name |
samples |
Array of strings | Yes | Sample identifiers in this group |
metadata |
Table | No | Arbitrary group-level metadata; each key is available as a wildcard in paths, shells, and when conditions (wildcard.<key>) |
Any rule that references {sample} or {group} is expanded once per (group, sample) pair across all groups.
Expanded rule naming: {rule_name}_{group}_{sample} (e.g., align_control_CTRL_001).
When groups are loaded from a TSV/CSV sheet, extra columns become metadata keys automatically — see Sheet columns as group metadata.
Loading groups from external file#
For large cohorts, use sample_groups_file in [workflow]:
TSV format (samples can be comma-separated within the field):
name samples reads_layout long_reads
short CTRL_001,CTRL_002 pe
long LR_001 ont reads/flowcell1.fq.gz;reads/flowcell2.fq.gz
Every row must supply a value for every column (tab-separated, one row per
line) — a row missing a trailing cell fails the sheet parse with an
unequal-lengths error. Leave a cell's value empty only if no when gate
or placeholder references that metadata key for the group's samples (an
empty cell behaves as an unbound key, which closes gates in both
directions — including !=).
JSON format:
[
{"name": "control", "samples": ["CTRL_001", "CTRL_002"]},
{"name": "case", "samples": ["CASE_001", "CASE_002"]}
]
Like pairs, a group name declared in both [[sample_groups]] and
sample_groups_file is deduplicated when identical (with a warning) and
rejected as a config error when the samples differ — see the merge rule in
Loading pairs from external file.
Sheet columns as group metadata#
In TSV/CSV sheets, every column other than name and samples becomes
group metadata: it fans out per (group, sample) pair as a bare
{key} wildcard in paths/shell/log and as wildcard.<key> in when
conditions — the same vocabulary as the inline
metadata table. This is how a cohort declares a second read
dimension (issue #283):
[[rules]]
name = "assemble"
input = ["{long_reads}"]
output = ["asm/{group}_{sample}/assembly.fa"]
when = "wildcard.reads_layout == 'ont' || wildcard.reads_layout == 'pacbio'"
shell = "flye --nano-raw {input} --out-dir {output[0]}"
[[rules]]
name = "qc"
input = ["raw/{sample}_R1.fastq.gz"]
output = ["qc/{sample}.txt"]
when = "wildcard.reads_layout != 'ont' && wildcard.reads_layout != 'pacbio'"
shell = "fastqc {input} -o {output[0]}"
Short-read rows gate assemble closed and qc open; long-read rows the
reverse — one rule template per branch, expanded per sample.
reads_layout vocabulary — the column is validated at plan time;
accepted values are pe, single, interleaved (short reads) and ont,
pacbio (long reads). Any other value fails the workflow with the accepted
set named, because a typo would silently close every when gate built on
it. Other metadata columns stay free-form — use the
metadata_file table for per-sample
free-form values.
Multi-file cells — a metadata cell may carry several paths (e.g. one
FASTQ per flowcell) separated by ; (or tab-separated list cells in JSON
sheets). They are normalized to a single space-joined value that splices
verbatim into the input list
as one element; shell rendering joins input-list elements, so
flye --nano-raw {input} receives all paths as separate shell tokens
(the oxo-flow-varlo target_regions multi-file pattern). Spaces alone
do not split cells — separate multiple paths with ;.
Reserved column names — group and sample cannot be used as sheet
columns: the fan-out binds {group}/{sample} first and metadata keys
last, so such a column would shadow the engine's binding for every row in
the group. A sheet carrying them is rejected at parse time — rename the
column (e.g. sample_id).
Example#
[[sample_groups]]
name = "treatment"
samples = ["S001", "S002"]
[[rules]]
name = "align"
input = ["raw/{sample}_R1.fq.gz"]
output = ["aligned/{sample}.bam"]
shell = "bwa mem ref.fa {input[0]} > {output[0]}"
Produces align_treatment_S001 and align_treatment_S002.
See examples/gallery/12_cohort_analysis.oxoflow for a complete cohort study pipeline.
Combining Groups and Pairs — duplicate sample ownership#
A sample may appear in both a user-declared [[sample_groups]] entry and an
[[pairs]] entry (e.g. every group sample is also paired against a
control). This is legal — group-wildcard rules and pair-wildcard rules
expand independently — but the engine warns at parse time, because unless
every rule path distinguishes the instances (e.g. a {group} component),
the instances share an output path and one is skipped as up-to-date while
consuming the other's data. The warning prints once per parse naming
all duplicates, not once per duplicate sample:
WARN oxo_flow_core::config::parse: 3 samples are declared by more than one owner: S1 (group 'case' / pair 'P1'), S2 (group 'case' / pair 'P2'), S3 (group 'case' / pair 'P3') — unless every rule path distinguishes them (e.g. a {group} component), the instances share an output path and one is skipped as up-to-date while consuming the other's data
A single duplicate renders a per-sample message instead:
sample 'S1' is declared by both group 'case' and pair 'P1' — {same note}.
Auto-discovered
samples (sample_pattern) feeding [[pairs]] are the documented pairing
workflow, not a competing owner — only user-declared groups are compared
against pair membership. Fix by adding a {group}/{pair_id} component to
the affected rules' output paths, or by keeping the domains disjoint.
Auto-Discovery with sample_pattern#
The [workflow] sample_pattern auto-discovers samples from the filesystem:
[workflow]
# Exactly one pattern — uncomment the one that matches your data:
# Paired-end reads (Illumina standard)
sample_pattern = "raw/{sample}_R1.fastq.gz"
# # Paired-end reads (common variant)
# sample_pattern = "raw/{sample}_1.fq.gz"
# # Single-end reads
# sample_pattern = "raw/{sample}.fastq.gz"
sample_pattern binds {sample} only — {config.*} placeholders are
expanded, but any other wildcard name ({read}, {replicate}, {lane}, …)
is captured by the glob and never bound: every matching file collapses to the
same {sample} value and expansion fails with duplicate rule name. Point
the pattern at one file per sample (R1 for paired-end data) and list the mate
in each rule's input:
sample_pattern = "raw/{sample}_R1.fastq.gz"
[[rules]]
name = "align"
input = ["raw/{sample}_R1.fastq.gz", "raw/{sample}_R2.fastq.gz"]
output = ["aligned/{sample}.bam"]
shell = "bwa mem ref.fa {input[0]} {input[1]} > {output[0]}"
Supported wildcards in sample_pattern:
| Wildcard | Description | Example match |
|---|---|---|
{sample} |
Sample identifier | SAMPLE_01 from SAMPLE_01_R1.fastq.gz |
For genuine per-read/per-replicate fan-out, declare the domain explicitly with
[[sample_groups]],
[[pairs]], or [[values]], and
reference the extra paths in each rule;
input_groups binds additional wildcards
discovered from files.
Merging Multiple Sample Sources#
All sample sources are merged into a single {config.samples_list}:
# Filesystem auto-discovery
sample_pattern = "raw/{sample}_R1.fastq.gz"
# CSV/TSV file
sample_groups_file = "metadata/samples.csv"
# Ad-hoc via CLI — filters the declared set (unknown names fail);
# declares samples only when the workflow ships none
oxo-flow run pipeline.oxoflow --samples EXTRA_01,EXTRA_02
All sources deduplicate — the same sample from multiple sources appears once.
Per-group sample lists are available as {config.samples_<group_name>}.
These injected values are comma-joined strings (e.g. "S001,S002"):
- In
expand_inputs,scatter.values_from, andsplit.values_fromthey resolve per value (comma-split) — see Expand Inputs. - In shell templates,
{config.samples_list}renders as the comma-joined text, e.g.for s in $(echo {config.samples_list} | tr ',' ' ').
Sample Metadata (metadata_file)#
Per-sample metadata is loaded at plan time from an external TSV/CSV/JSON
table and addressed as {meta.<column>} in any rule's input, output,
shell, log, script, hooks, and when. This is the
metadata group/pair-key
mechanism generalized to an arbitrary per-sample column vocabulary — the
methylseq when = "config.single_end_mode || {meta.endedness} == 'SE'"
pattern, chipseq antibody fan-out (see
Input Groups), and rnaseq-style per-sample
parameter lookups:
TSV/CSV format — the first column is the sample id (must match the
values your {sample} wildcards resolve to); every other header becomes a
column:
sample endedness adapters
SE1 SE AGATCGGAAGAGCACACGTCTGAACTCCAGTCA
PE1 PE AGATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT
JSON format — a list of objects; the sample key identifies the row
(all other keys become columns):
[
{"sample": "SE1", "endedness": "SE", "adapters": "AGATCGGAAGAGCACACGTCTGAACTCCAGTCA"},
{"sample": "PE1", "endedness": "PE", "adapters": "AGATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT"}
]
Semantics:
- The table is a lookup only — it does not define the sample set. Use
sample_groups/sample_groups_file/sample_patternfor that. - At expansion, each rule instance resolves its metadata row from its
{sample}binding, falling back to{pair_id}, then{experiment}, then{control}(first binding that has a row wins). - A missing row or column renders empty in shells/inputs/outputs, with
a plan-time warning for a column that no row defines (the same
stance as
{values.name}). - In
when, metadata values are baked as quoted literals in comparisons ({meta.endedness} == 'SE'→'SE' == 'SE') andtrue/falsein truthiness positions; a missing row or column renders''/false— the gate is closed, so samples without the data evaluate false instead of running. - A
{meta.<column>}reference in a rule that never fans out (no sample-like binding) is left as a residual placeholder — the engine logs a warning at execution time and the literal text remains in the rendered command.
[[rules]]
name = "trim"
input = ["raw/{sample}_R1.fastq.gz"]
output = ["trimmed/{sample}.fq"]
when = "config.single_end_mode || {meta.endedness} == 'SE'"
shell = "cutadapt -a {meta.adapters} {input[0]} -o {output[0]}"
Metadata paths can also feed inputs directly (DAG edges still infer by exact path match, since the paths are plan-time literals):
[[rules]]
name = "mapping_qc"
input = ["{meta.bam_path}"]
shell = "samtools stats {input[0]} > {output[0]}"
Partial Pair Tolerance & Tumor-Only Mode#
Pairs with missing controls are supported — control is now optional:
# Tumor-only CNV calling (no matched normals)
[[pairs]]
pair_id = "T1"
experiment = "T1"
# control omitted → {control} = ""
# Pooled normal via config values
[[pairs]]
pair_id = "T2"
experiment = "T2"
# control = "" # empty string also works
Rules using {control} receive an empty string when no control is specified.
Use shell conditionals or declarative [config] entries to handle this:
[[rules]]
name = "cnv_detect"
input = ["aligned/{experiment}.bam"]
shell = """
if [ -n "{control}" ]; then
cnvkit.py batch {input} --normal aligned/{control}.bam
else
cnvkit.py batch {input} --method cbs # tumor-only mode
fi
"""
For pooled-normal scenarios, use declarative [config] entries:
[config]
normal_mode = { default = "pooled", help = "matched, pooled, or none" }
pooled_normal = { default = "results/pooled_normal.bam" }
Supported multi-omics pair patterns:
| Scenario | Pair Configuration | Control |
|---|---|---|
| Matched tumor-normal | experiment = "T1", control = "N1" |
Required |
| Unmatched tumor vs pooled | experiment = "T1" |
None |
| Tumor-only (CNV, somatic) | experiment = "T1" |
None |
| Paired-end case-control | experiment = "CASE", control = "CTRL" |
Required |
| Time-series (no control) | two entries: experiment = "T0" and experiment = "T6" |
None |
when — Conditional Rule Execution (WF-01)#
The optional when field on a rule contains an expression evaluated against [config] values. When the expression evaluates to false the rule is skipped at execution time (JobStatus::Skipped, "condition evaluated to false") — it remains in the DAG but does not run. Predicates over per-instance bindings (wildcard.<key>, {meta.<column>}) are decided earlier, at plan time: non-matching instances never enter the DAG, so dry-run and validate counts reflect exactly what a real run would execute.
[[rules]]
name = "fastqc"
when = "config.run_qc"
input = ["raw/sample_R1.fq.gz"]
output = ["qc/sample_fastqc.html"]
shell = "fastqc {input[0]} -o qc/"
Expression syntax#
| Form | Example | Description |
|---|---|---|
config.<key> |
config.run_qc |
Truthy check (true, non-zero, non-empty string) |
config.<key> == "value" |
config.mode == "WGS" |
String equality |
config.<key> != "value" |
config.mode != "WES" |
String inequality |
config.<key> == true\|false |
config.skip == false |
Boolean equality |
config.<key> > N |
config.min_cov >= 20 |
Numeric comparison (>, >=, <, <=) |
len(config.<key>) |
len(config.gene_sets) > 0 |
Element count comparison: array items, string characters, or table keys, compared against a whole number (==, !=, >=, <=, >, <). A missing key counts as 0. Bare len(config.<key>) is truthy when non-empty. Note a bare config.<key> array is truthy even when empty — len(...) > 0 is the "list is non-empty" gate. Comparing against a non-numeric literal is a lint warning (W029) and always evaluates false |
file_exists("path") |
file_exists("panel.bed") |
File existence test. Relative paths resolve against the workflow root (the workflow file's directory), not the process working directory, and every evaluation site (plan time, execution re-check, config-change analysis) resolves the same way. In per-instance predicates (wildcard.<key> / {meta.<column>}) it is evaluated at plan time — the instance set is fixed when the DAG is built, so files produced mid-run cannot gate instances; express those with {meta.*} (plan-time bakeable) or an execution-time config gate instead |
reads_count("path") / wc_lines("path") / file_size("path") |
reads_count('trimmed/{sample}.fq.gz') > config.min_reads |
Data-dependent gates — count FASTQ records (4-line records, plain or .gz), text lines of the decompressed content (plain or .gz), or file bytes in the run workdir at execution time. See Data-Dependent when Gates |
regex_extract("path", "pattern", group?) |
regex_extract('alignment/{sample}.flagstat', '(\\d+) \\+ \\d+ mapped', 1) > config.min_mapped_reads |
Data-dependent gate — parse an integer out of a text report's first regex match (issue #312). group defaults to 0 (the whole match — the pattern must then match a bare number); use group 1+ to extract a number embedded in surrounding text (samtools flagstat's mapped count, bcftools-stats record counts). The captured text must parse as an integer; missing file / no match / non-numeric capture fail closed at execution time. Malformed calls (wrong arity, non-numeric group, uncompilable pattern) are flagged as W030 |
wildcard.<key> == "value" |
wildcard.control != "" |
Per-instance pair/group wildcard comparison (pair_id, experiment, control, tumor, normal, experiment_type, tumor_type, group, sample, plus any metadata keys declared on [[pairs]] / [[sample_groups]] and [[values]] table names) — evaluated per expansion combo at DAG build time; non-matching instances never enter the DAG (snakemake-style morphing) |
{meta.<column>} == "value" |
{meta.endedness} == 'SE' |
Per-instance sample-metadata comparison (see Sample Metadata) — baked and evaluated per instance at plan time; instances whose gate evaluates false never enter the DAG. A missing row or column renders '' in a comparison (a closed gate) |
!<expr> |
!config.skip |
Logical NOT |
<expr> && <expr> |
config.run_qc && config.min_cov >= 20 |
Logical AND |
<expr> \|\| <expr> |
config.wgs \|\| config.wes |
Logical OR |
(<expr>) |
(config.a && config.b) \|\| config.c |
Grouping |
Any other token — a bare identifier like params_mode, or an unknown function call — is a validation error (E018): the evaluator treats unrecognizable atoms as true, so a typo'd condition would silently let the rule run. Reference config keys as config.<key>, or use a quoted literal, a {...} placeholder, or one of the built-in functions.
Data-Dependent when Gates#
The runtime functions reads_count(path), wc_lines(path),
file_size(path), and regex_extract(path, pattern, group?) let a when
gate inspect files produced earlier in the same run (issues #282/#312) —
e.g. skip a variant caller on samples whose trimmed FASTQ has too few
records:
[[rules]]
name = "call_variants"
when = "reads_count('trimmed/{sample}.fq.gz') > config.min_reads"
input = ["trimmed/{sample}.fq.gz"]
output = ["variants/{sample}.vcf.gz"]
shell = "caller --input {input[0]} --output {output[0]}"
- Path resolution — relative paths join against the run workdir;
absolute paths pass through.
{sample}and other wildcards expand from the instance context before the read. When--workdirdiffers from the workflow directory, workflow-shipped files are anchored automatically: relative env file specs (conda/mamba/pixi/venv_requirements), rule scripts,input_groupsscan patterns, and[[references]].sourceresolve against the workflow directory.{input}/{log}in rule commands stay workdir-relative — data and logs are addressed from the run cwd (matching the freshness gates and the engine's log-dir creation), so point--workdirat the directory holding your data (issue #427). - Phase semantics — at plan time a runtime-function atom whose file
is missing — or whose file is an output of this run but still holds a
previous run's stale value — defers (
true, the rule stays in the plan because its producer may not have run yet); at execution time a missing or unreadable file fails closed (false, the rule is skipped). The defer propagates through!,&&,||and parentheses (three-valued logic), so a negated or chained gate is never pruned at plan time just because its target file is not ready. Planning withoxo-flow dry-runtherefore never prunes on a gate whose target file does not exist yet. - Formats —
reads_countcounts 4-line FASTQ records in plain or gzip files (the record count is truncated — a trailing partial record does not round up);wc_linesstreams lines of the decompressed content for plain or gzip files (same.gzdetection asreads_count);file_sizereturns the byte length. BAM/CRAM input to these functions is out of scope for now (it needs a BGZF parser, not line arithmetic). - Regex extraction —
regex_extract(path, pattern, group?)reads the whole file (plain or.gz, up to 16 MiB) and takes the first regex match; the captured text (group0= whole match by default, or an explicit capture group) must parse as an integer. Matching is line-oriented ((?m)is applied automatically), so^and$anchor per line. Commas inside the pattern's{}quantifiers,[]character classes, or quoted arguments do not split arguments. The extracted value compares against literals andconfig.*thresholds exactly like the counting functions. - Change tracking — recorded verdicts participate in config-change
replay: changing a threshold referenced by such a gate (e.g.
config.min_reads) re-evaluates only the instances whose verdict actually flipped, keeping completed unchanged instances exempt. - Checkpoint accounting — verdicts are stored in the checkpoint's
when_verdictsmap. A gate that evaluated tofalseis a skip, not a failure: the rule never entersfailed_rules(a stale entry is pruned),statusreports it under Skipped (when condition false) (skipped_by_whenin--json), andreport --r-datawrites it tometrics.tsvasskipped_by_when(issue #690). - Scope — only
whenstrings may call these functions;shell,input, andoutputcannot read the filesystem at plan time.
Unbound wildcard keys evaluate false#
A wildcard.<key> predicate whose key has no binding for an instance —
e.g. a group that declares no input_type metadata while another group
does — evaluates to false for every operator (==, !=, >, <,
…, and bare truthiness alike). The condition is not met, so that instance
never runs. This keeps a rule like
when = "wildcard.input_type == 'srr'" scoped to the SRA cohort even when
other cohorts omit the key entirely.
Note: this is a behavioral contract, not an oversight. An unbound comparison historically evaluated
true, which ran SRA-download rules against literal{accession}placeholders for FASTQ cohorts (live incident).oxo-flow lintflagswhenreferences to keys no pair/group can ever bind as a W027 warning — such rules never run.
Per-instance wildcard and {meta.<column>} bindings are baked into a kept
rule's when at expansion time, so the execution-time re-check
re-evaluates the exact same per-instance verdict; config-only predicates
continue to gate at execution as before.
wildcard.* is a when-only vocabulary. The dotted form is not a
shell placeholder — the engine's placeholder syntax is bare {name}, so
{wildcard.input_type} in a shell renders literally. Group/pair
metadata keys reach shells as bare placeholders ({input_type}), baked
per instance at plan time when the rule fans out on {group}/{sample};
a rule that does NOT fan out per sample has no binding to bake, so such
keys stay literal there (use input_groups or {meta.<column>} for
those cases).
{meta.<column>} truthiness and boolean columns#
In a bare truthiness position (when = "{meta.do_qc}") a metadata value
bakes to true/false with shell-style truthiness: non-empty and not
false and not 0 is true — so "0", "", and "false" close the
gate and everything else (including "foo") opens it. For a strict
boolean column, use the comparison form instead so any unexpected value
falls through to no override rather than opening both branches of a
markdup/nomarkdup pair:
# strict per-sample boolean (recommended for true/false columns)
when = "config.mark_duplicates || {meta.mark_duplicates} == 'true'"
A missing row or column renders false in a truthiness position and ''
in a comparison — always a closed gate, never an unbound token. This makes
{meta.x} == '' the official absence guard idiom: it gates the
"column present" branch and closes cleanly for rows that lack the column.
Example#
[config]
run_annotation = true
min_coverage = 30
mode = "WGS"
[[rules]]
name = "vep_annotate"
when = 'config.run_annotation && config.min_coverage >= 20'
# ...
[[rules]]
name = "wgs_coverage"
when = 'config.mode == "WGS"'
# ...
See examples/gallery/11_conditional_workflow.oxoflow for a full example.
Dependency Resolution#
Dependencies are inferred automatically: if rule B lists a file in its input that appears in rule A's output, then B depends on A.
[[rules]]
name = "step1"
output = ["intermediate.txt"]
# ...
[[rules]]
name = "step2"
input = ["intermediate.txt"] # depends on step1
# ...
No explicit dependency declaration is needed.
transform — Unified Scatter-Gather Operator#
The transform operator unifies split → map → combine patterns into a single rule declaration, similar to dplyr's group_by() %>% summarize() or pandas' groupby().apply().
Structure#
[[rules]]
name = "variant_calling"
input = ["aligned/sample.bam"]
output = ["variants/sample.vcf.gz"]
[rules.transform.split]
by = "chr"
values_from = "config.chromosomes"
[rules.transform]
map = "gatk HaplotypeCaller -R {config.reference} -I {input} -L {chr} -O {output}"
cleanup = true
[rules.transform.combine]
shell = "gatk GatherVcfs $(for f in {chunks}; do echo \"-I $f \"; done) -O {output}"
Split Configuration#
| Field | Type | Description |
|---|---|---|
by |
String | Required. Variable name for splitting (e.g., "chr", "sample") |
values |
Array | Direct list of split values |
values_from |
String | Reference to config variable (e.g., "config.chromosomes") |
n |
String | Number of map instances — the split values are the labels "0", "1", ..., "n-1". It does not partition the input data |
glob |
String | Glob pattern to find split values from files |
Priority: values → values_from → n → glob
Split values are labels, not data partitions
The engine never reads, slices, or chunks the declared input. It repeats
the rule once per split value and substitutes that value wherever
{split_var} appears. With n = "3" and input = ["in.txt"], the map
command runs three times with the same input:
process in.txt > .oxo-flow/chunks/chunk/0.txt
process in.txt > .oxo-flow/chunks/chunk/1.txt
process in.txt > .oxo-flow/chunks/chunk/2.txt
Dividing the work is the map command's job: it must use the split value
to select its own slice (e.g. -L {chr} for per-chromosome calling), or
the rule's input must be the split value (input = ["{chr}"]), so
each instance reads its own file. The engine never slices a file, and a
split variable embedded in a longer path (shard_{chr}.txt) is left
literal. n = "3" alone does neither.
Combine Configuration#
| Field | Type | Description |
|---|---|---|
shell |
String | Shell command to combine chunks |
aggregate |
Boolean | Enable automatic aggregation |
method |
String | Aggregation method: "concat" or "json_merge" |
header |
String | Header line for concat aggregation |
Built-in Variables#
| Variable | Expands to |
|---|---|
{split_var} |
Current split value (e.g., {chr} → "chr1") |
{chunks} |
Space-separated list of all chunk outputs |
{input} |
Chunk outputs — same as {chunks} (in combine) |
{output} |
Original rule output (in combine) |
Modes#
Mode A: Split → Map → Combine
Classic scatter-gather with explicit combine command:
[rules.transform.split]
by = "chr"
values_from = "config.chromosomes"
[rules.transform]
map = "gatk HaplotypeCaller -R {config.reference} -I {input} -L {chr} -O {output}"
[rules.transform.combine]
shell = "gatk GatherVcfs $(for f in {chunks}; do echo \"-I $f \"; done) -O {output}"
Mode B: Split → Map (No Combine)
Parallel processing without merging — each split produces independent output:
[rules.transform.split]
by = "chr"
values_from = "config.chromosomes"
[rules.transform]
map = "samtools flagstat {input} > {output}"
# No combine section
Mode C: Split → Map → Aggregate
Automatic aggregation (concat or json_merge):
[rules.transform.split]
by = "chunk"
n = "5"
[rules.transform]
map = "process {input} > {output}"
[rules.transform.combine]
aggregate = true
method = "concat"
Cleanup#
When cleanup = true, chunk files are automatically deleted after the whole run finishes successfully (empty chunk directories are removed too):
Failed runs keep their chunks for debugging. Because chunk deletion makes the map outputs "missing", a re-run always recomputes the map rules — that is the disk-space trade-off of cleanup = true.
Failure and Retry Logic#
In a scatter-gather process, failures are handled at the chunk (map) level:
- If a single chunk fails, only that specific chunk is retried according to the rule's
retriessetting. - Sibling chunks continue to process in parallel.
- The combine step will not execute until all chunks succeed. If any chunk fails exhaustively (after all retries), the combine step is cancelled.
Expanded Rule Naming#
Transform rules expand into:
- Map rules:
{rule_name}_{split_value}(e.g.,variant_calling_chr1) - Combine rule:
{rule_name}_combine(e.g.,variant_calling_combine)
Multi-line Strings#
Use triple quotes for multi-line shell commands:
shell = """
mkdir -p results
bwa mem -t {threads} ref.fa {input} | \
samtools sort -@ {threads} -o {output}
"""
Backslashes inside \"\"\" blocks are TOML escape sequences
"""…""" is a TOML basic string: escape sequences like \n, \t, and
\\ are decoded before the shell ever sees the text. Two consequences:
- The
\line continuations above only work because TOML folds backslash-newline into nothing. But an embedded heredoc or regex that needs a literal\nbreaks — the shell receives a real newline instead. - For scripts with literal backslashes (embedded
python3 <<'EOF'heredocs,sed 's/\t/,/g', …), use a multi-line literal string with triple single quotes — TOML passes its contents through byte-for-byte:
Complete Example#
[workflow]
name = "ngs-pipeline"
version = "2.0.0"
description = "Complete NGS analysis pipeline"
author = "Shixiang Wang <w_shixiang@163.com>"
[config]
reference = "/data/ref/hg38.fa"
known_sites = "/data/ref/known_sites.vcf.gz"
results = "results"
[defaults]
threads = 4
memory = "8G"
environment = { conda = "envs/base.yaml" }
[report]
format = ["html"]
[[rules]]
name = "fastqc"
input = ["raw/{sample}_R1.fastq.gz", "raw/{sample}_R2.fastq.gz"]
output = ["{config.results}/qc/{sample}_R1_fastqc.html"]
shell = "fastqc {input} -o {config.results}/qc/ -t {threads}"
[[rules]]
name = "trim"
input = ["raw/{sample}_R1.fastq.gz", "raw/{sample}_R2.fastq.gz"]
output = [
"{config.results}/trimmed/{sample}_R1.fastq.gz",
"{config.results}/trimmed/{sample}_R2.fastq.gz"
]
environment = { docker = "biocontainers/fastp:0.23.4" }
shell = "fastp --in1 {input[0]} --in2 {input[1]} --out1 {output[0]} --out2 {output[1]} --thread {threads}"
[[rules]]
name = "align"
input = [
"{config.results}/trimmed/{sample}_R1.fastq.gz",
"{config.results}/trimmed/{sample}_R2.fastq.gz"
]
output = ["{config.results}/aligned/{sample}.bam"]
environment = { conda = "envs/alignment.yaml" }
shell = "bwa mem -t {threads} -R '@RG\\tID:{sample}\\tSM:{sample}' {config.reference} {input[0]} {input[1]} | samtools sort -o {output}"
[rules.resources]
threads = 16
memory = "32G"
JSON Schema#
oxo-flow ships a JSON Schema for the .oxoflow format. Use it for automated validation in your CI/CD pipelines or for real-time autocompletion and error checking in your IDE (like VS Code or IntelliJ).
Getting the Schema#
You can output the schema directly from the CLI:
IDE Configuration (VS Code)#
To enable validation in VS Code, add the following to your settings.json:
"yaml.schemas": {
"https://traitome.github.io/oxo-flow/schema/oxoflow-v1.schema.json": "*.oxoflow"
}
(Note: Although .oxoflow is TOML, many VS Code extensions can apply JSON schemas to multiple formats).
Checkpoint Re-entry#
A checkpoint = true rule discovers new values at runtime (Snakemake-style
checkpoint re-entry): after it completes, the engine reads its
checkpoint_manifest — a TOML file the rule itself wrote — and re-expands
the workflow with the new values. Every round is still a static plan, so
previews stay deterministic and resumes reconstruct the same plan.
[[rules]]
name = "discover"
output = ["discover.toml"]
shell = "python discover.py > discover.toml"
checkpoint = true
checkpoint_manifest = "discover.toml"
The manifest declares new wildcard values — new samples and/or new experiment-control pairs in the same round:
[reentry]
group = "batch" # optional; defaults to the workflow's first sample group
sample = ["S4", "S5"] # appended (dedup) to that group
pairs = [ # optional; appended (dedup by pair_id) to [[pairs]]
{ pair_id = "CASE_007", tumor = "T7", normal = "N7", tumor_type = "tumor" },
]
Pair entries mirror [[pairs]] fields: pair_id (required), experiment
(required, alias tumor), control (alias normal), experiment_type
(alias tumor_type), and an optional metadata table. tumor_type and
metadata are optional.
Semantics:
- On success, the engine merges the samples and pairs, then re-expands the
rule templates; newly created instances (e.g.
analyze_batch_S4,call_CASE_007) execute in the same run. Existing instances are untouched. - The checkpoint records each re-entry (
reentriesarray: round, rule, group, samples, pairs). A resume replays the records whose checkpoint rule is still up-to-date — the plan reconstructs identically. If the checkpoint rule is invalidated (input/config change), its contribution is revoked (samples and pairs) until it re-runs and re-records. - Pair identity is the
pair_id: an existing id with identical content is a no-op; an existing id with different content is an error (E015) — silently superseding it would corrupt already-run pair outputs. The same sample appearing in several pairs is not an ambiguity: pair instances are keyed bypair_id, so each pair is its own instance. - A discovered
pair_idwhose instance name collides with an existing instance (e.g. a sample-group instanceanalyze_CASE_007) is an error (E016). - An empty manifest (
sample = [], nopairs) is a valid no-op. A missing or unparsable manifest fails the checkpoint rule with a clear error, and its dependents do not run. - Bounds: checkpoint rules are never re-expanded themselves (no
{sample}/{group}/{pair_id}/{experiment}— validation error E014) and re-entry is capped at 32 rounds — a rule that keeps discovering values past that is a workflow bug, not an engine feature. Diagnostic E013 (emitted at warning severity —checkpoint = truepredates the manifest field) flags a checkpoint rule that does not declarecheckpoint_manifest. dry-runpreviews replay recorded re-entries (the preview shows the same static plan a run would execute) and mark checkpoint rules as possible re-entry points;--jsonincludes areentrysection.
Runtime-Discovered Outputs (output_pattern)#
An output_pattern rule declares its outputs by filesystem pattern
instead of a fixed output list: after the rule completes, the engine scans
the working directory for files matching the pattern, and any downstream
rule whose input references the pattern's wildcard is instantiated once per
discovered value. This is the "checkpoint-lite" variant of re-entry — the
discovered values come from the file system, not from a manifest the rule
wrote (which is what makes it work for rules whose file counts are unknown
ahead of time, e.g. interval scatters or contig-aware variant calling).
[[rules]]
name = "split"
output_pattern = "results/chunks/{i}.txt"
shell = "for i in 1 2 3; do echo $i > results/chunks/$i.txt; done"
[[rules]]
name = "collect" # deferred: no {i} instances at plan time
input = ["results/chunks/{i}.txt"]
output = ["results/merged/{i}.txt"]
shell = "cat {input} > {output}"
split produces results/chunks/1.txt, results/chunks/2.txt,
results/chunks/3.txt; the engine instantiates collect_1, collect_2,
collect_3 at runtime and inserts them into the plan, so the run completes
with all three merged files. The fresh wildcard may combine with existing
sample wildcards — a producer fanned out over {sample} contributes its
domain per sample instance, and consumers instantiate over the union.
Semantics:
output_patternis mutually exclusive withoutputand withtransformandinput_groups(validation errors; v1 exclusions). The pattern needs at least one wildcard — a fixed pattern can never enumerate a domain.{config.x}in the pattern resolves from the workflow config before the scan; discovered files are literals, so exact-match edge inference works as usual.- A
[[values]]wildcard in the pattern is not fresh — it is a plan-time bound dimension (issue #296). The producer fans out per value with the pattern baked per instance, the pattern is registered producer-side in the DAG (so a per-value consumer's materialized input exact-matches its own producer instance), and a pattern whose wildcards are all bound this way skips the runtime scan entirely — the static instances already are the domain. Mixed patterns (results/{assembler}/{chunk}.txt) compose: the values dimension bakes per instance while the fresh wildcard stays runtime-discovered, and the deferred consumers instantiate over the union with both bindings. - Plan time: a consumer referencing a fresh wildcard has an empty domain and
instantiates nothing; it is deferred whole. Wildcards shared with the
producer's pattern ride the discovered bindings;
[[values]]tables the consumer references on its own project as an extra Cartesian dimension at instantiation, under the same plan-time semantics as a static fan-out —wildcard_constraintsfiltered and per-instancewhengates evaluated, so gated-off combos never become instances (issue #296 follow-up). A consumer declared before its producer is legal but warned (output_pattern consumer declared BEFORE its producer; the fan-out still resolves, but declare the producer first for stable run attribution). Referencing two producers' fresh wildcards in one rule is a v1 error; a chain — a rule that is both consumer and producer of fresh wildcards — is supported. - Runtime: when a producer instance completes, the engine scans its
pattern and unions the discovered values into the producer template's
domain. When all of the producer's instances have completed, the
deferred consumers are instantiated (one instance per domain value × own
values binding, deterministically named
consumer_<own values>_<values>, withexpand_inputspatterns materialized into concrete inputs), and the new instances are inserted forward into the remaining plan — forward-safety comes from completion ordering, not DAG topology, exactly like checkpoint re-entry. - Interaction with
resume: the discovered domains are persisted (persist-first, before fan-out). A resume replays the persisted domain and instantiates the deferred consumers only when every producer instance is already completed; when a failed instance still has to re-run, its consumers stay pending and the live runtime pass instantiates them from the final domain union — the final consumer set equals the union of all per-instance discoveries, never a partial slice. - Zero discoveries after producer success is loud: a warning is emitted and the consumers stay uninstantiated (a later producer instance may still contribute). A producer failure leaves the consumers uninstantiated and the failure summary attributes them — "not instantiated: producer X failed" — never a silent zero-match.
- Checkpoints persist the discovered domain (
output_pattern_domains), so aresumere-instantiates the consumers without re-running the producer; re-instantiation is idempotent (same names, same plan).
[webhook] — Workflow-Level Notifications#
Post workflow_started / workflow_completed / workflow_failed events
to an HTTP endpoint when a run starts and finishes (issue #227 item 1) —
the transport counterpart of the on_complete / on_error shell hooks.
[webhook]
url = "https://hooks.example.com/oxo"
# method = "POST" # POST (default) | PUT | GET
# events = ["workflow_completed", "workflow_failed"] # default: workflow_completed
# headers = { Authorization = "Bearer …" }
# secret = "…HMAC key…" # X-OxoFlow-Signature header (HMAC-SHA256)
# signature_scheme = "hmac-sha256"
# timeout_secs = 30
# max_retries = 3
| Field | Type | Default | Description |
|---|---|---|---|
url |
String | — (required) | Webhook endpoint URL |
method |
String | POST |
HTTP method (POST/PUT/GET) |
events |
Array of String | ["workflow_completed"] |
Which events fire: workflow_started, workflow_completed, workflow_failed (rule_completed/rule_failed are defined but not yet dispatched) |
headers |
Table | {} |
Custom request headers |
secret |
String | — | HMAC key for the X-OxoFlow-Signature header |
signature_scheme |
String | hmac-sha256 |
hmac-sha256 (RFC 2104) or the legacy sha256-keyed |
timeout_secs |
Integer | 30 |
Per-request timeout |
max_retries |
Integer | 3 |
Retries with exponential backoff |
The payload is JSON: {event, workflow_name, timestamp, data, version}
where data carries the run counters (succeeded, failed, skipped)
and, on failure, an error summary. Notifications are best-effort —
an unreachable or erroring endpoint logs a warning and never changes the
run status or its exit code.
[engine] — Runtime Disk-Pressure Guard (issue #843)#
Optional run-time disk monitoring. Without this section the engine never measures free space mid-run and behavior is unchanged.
[engine]
min_free_disk = "50G" # required to enable: abort/park floor
# reclaim_free_disk = "100G" # optional: reclaim trigger; defaults to min_free_disk
| Field | Type | Default | Description |
|---|---|---|---|
min_free_disk |
String | — | Floor for free disk space at the working directory. A bare number means MiB. Sizes accept T/G/M/K suffixes (binary units: 1G = 1024 MiB). Note "50GB" is not valid — the unit is a single letter |
reclaim_free_disk |
String | min_free_disk |
Reclaim trigger level; must be ≥ min_free_disk (validation error otherwise) |
Response ladder#
Each measurement (run start, loop head, and every 60 s while the run is parked or between dispatch batches) classifies the working directory's free space and reacts:
| Level | Condition | Response |
|---|---|---|
| Normal | free ≥ reclaim |
Run proceeds normally |
| Reclaim | min ≤ free < reclaim |
Completed rules' temporary outputs are reclaimed (tombstoned in the checkpoint + unlinked) to make room; the run continues |
| Hold | free < min |
New rule dispatches are withheld (in-flight rules continue); if nothing is in flight and nothing is reclaimable, the run parks and re-measures every 60 s |
| Abort | Hold persists for ~2 consecutive parked rounds (~2 min) | The run checkpoints its state and fails with the disk-pressure message |
Under Hold with rules in flight, in-flight rules run to completion (their outputs may themselves trigger Reclaim), but no new rule starts until free space recovers above the floor.
The abort message explains recovery: free space manually (delete
intermediates or run oxo-flow clean), then oxo-flow resume from the
checkpoint — or lower min_free_disk.
What gets reclaimed#
Reclamation targets a completed rule's outputs only when all of the
following hold: the rule is marked temporary = true, it is not a leaf (it
has dependents), every dependent has completed and each dependent's own
declared outputs still exist on disk — a completed rule whose outputs went
missing will re-execute and needs its inputs in place, so reclaiming them
underneath it would break the re-run (the reclaim-vs-resume race, fixed
under issue #843), the paths are not protected_output, and the files
exist on disk. Reclaimed rules get a tombstone in checkpoint.json — the
same mechanism the success-path temporary cleanup uses — so a later run
that needs the outputs regenerates them first (lazy cascade-up). See
Cleanup Behavior.
Notes:
- Measurements are fail-open: if the working directory cannot be stat'ed, the engine logs a warning and continues rather than aborting.
- On non-Unix platforms the monitor is inactive (a warning is printed at
startup when
[engine]thresholds are set). - A run that starts already below
min_free_diskprints a warning before executing (it does not refuse to start; the Hold/Abort ladder governs from there). - The park/abort window is tunable via environment:
OXO_FLOW_DISK_TICK_SECS(seconds between measurements while parked, default 60) andOXO_FLOW_DISK_HOLD_ROUNDS(grace park rounds before the abort, default 2). Unset, invalid, or0values fall back to the defaults, so a typo can never disable the abort. - Free space is measured at the working directory's filesystem (statvfs),
which is volume-level: thresholds apply to the whole volume, not to a
per-directory quota. Tools that spill large scratch outside the
workdir (honor
TMPDIRdeliberately if needed) are invisible to the monitor. - Mid-run reclaim of fan-out producers (
output_patternwith per-instance values) does not fire: the event loop tracks reclaim candidates by rule name and holds no instance bindings. Fail-safe — no reclaim rather than a wrong one. The post-run sweep covers them (see below). - Only paths declared by the rule (its
outputentries / matchingoutput_patterninstances) are reclaimed; sidecar files a command writes without declaring are never touched. Declare index sidecars explicitly (e.g. agenome.farule also declaringgenome.fa.fai) so reclaim + cleanup can track them.
See Also#
- Create a Workflow — practical authoring guide
- DAG Engine — how dependencies are resolved
- Environment System — environment specification details