Skip to content

Workflow Format#

The .oxoflow file format is oxo-flow's TOML-based workflow definition language. This page provides the complete specification, design philosophy, and syntax rules.


Design Principles#

The .oxoflow format is built on four core principles:

  1. Declarative over Imperative — Define what should happen (inputs, outputs, tools), not how to orchestrate it. The engine handles the execution logic.
  2. Explicit is better than Implicit — Every dependency and environment should be clearly visible. No hidden global state.
  3. Composition over Inheritance — Reuse logic through modular include directives and rule templates rather than complex inheritance hierarchies.
  4. Traceability by Default — The format structure directly supports generating provenance and audit trails (workflow checksums, per-rule input manifests, output hashes).

TOML Primer#

oxo-flow uses the TOML (Tom's Obvious, Minimal Language) format. If you are new to TOML, here are the three essential concepts used in .oxoflow files:

  1. Key-Value Pairs: key = "value". Strings must be in quotes.
  2. Tables: [name] defines a section (an object/map).
  3. Arrays of Tables: [[name]] defines a list of sections. In oxo-flow, rules are defined using double brackets because a workflow contains multiple rules.

For more details, see the Official TOML Specification.


File Extension#

Workflow files must use the .oxoflow extension (e.g., qc_pipeline.oxoflow).


Top-level Structure#

reference_dir = "..."   # Optional: base directory for auto-derived reference paths
[workflow]          # Required: metadata
[config]            # Optional: user variables (plain or declarative form)
[defaults]          # Optional: rule defaults
[report]            # Optional: report configuration
[[include]]         # Optional: include external workflow files
[[references]]      # Optional: auto-built indexes and reference data
[[rules]]           # Required: one or more rules
[[pairs]]           # Optional: experiment-control pairs (WC-01)
[[sample_groups]]   # Optional: multi-sample groups (WC-02)
[[values]]          # Optional: named value lists for parameter wildcards
[metadata]          # Optional: per-sample metadata columns ({meta.<column>})
[resource_budget]   # Optional: resource limits (enforced on cluster submissions; local runs use `--max-threads`/`--max-memory`)
[env_groups]        # Optional: named reusable environment specs
[resource_groups]   # Optional: shared resource pools (API limits, DB connections)
[wildcard_constraints] # Optional: regex patterns to constrain wildcard values
[cluster]           # Optional: HPC cluster profile (SLURM, PBS, SGE, LSF)
[[reference_db]]    # Optional: tracked reference database versions
[citation]          # Optional: citation metadata (DOI, authors, etc.)
[plugins]           # Optional: plugin configuration
[webhook]           # Optional: workflow-level webhook notifications (issue #227)
[engine]            # Optional: runtime disk-pressure guard (issue #843)

reference_dir is a bare top-level key, so it must appear before the first [table] header; it is also accepted inside [config]. See Auto-Derivation from reference_dir.


[[include]] — Modular Workflow Composition#

Include external workflow files to enable modular, reusable workflow design:

[[include]]
path = "common/qc.oxoflow"
namespace = "qc"

[[include]]
path = "align.oxoflow"
Field Type Required Description
path String Yes Path to the included .oxoflow file
namespace String No Optional namespace prefix for included rule names

Namespace Behavior#

When a namespace is specified:

  1. All rule names from the included file are prefixed: namespace::rule_name
  2. Internal depends_on references within the included file are automatically prefixed
  3. External depends_on references (to rules outside the included file) remain unchanged

Example:

# qc.oxoflow
[[rules]]
name = "fastqc"
input = ["{sample}.fastq.gz"]
output = ["qc/{sample}_fastqc.html"]
shell = "fastqc {input}"

[[rules]]
name = "trim"
input = ["{sample}.fastq.gz"]
output = ["trimmed/{sample}.fastq.gz"]
depends_on = ["fastqc"]  # Internal reference - will become "qc::fastqc"
shell = "fastp {input} -o {output}"
# main.oxoflow
[[include]]
path = "qc.oxoflow"
namespace = "qc"

[[rules]]
name = "align"
input = ["trimmed/{sample}.fastq.gz"]
depends_on = ["qc::trim"]  # Reference to included rule with namespace
shell = "bwa mem ref.fa {input} > aligned/{sample}.bam"

Resulting rules: qc::fastqc, qc::trim, align

Interface Contracts#

An [[include]] may declare a typed interface — what the host must wire in, what the module exposes, and config defaults it reads (issue #112). All checks run at validation time; execution semantics are unchanged, and includes without a contract behave exactly as before.

# qc.oxoflow
[[rules]]
name = "fastqc"
input = ["raw/{sample}.fq.gz"]      # the module's NEED
output = ["qc/{sample}_fastqc.html"]
shell = "fastqc {input} -o {output}"

[[rules]]
name = "internal"
input = ["qc/{sample}_fastqc.html"]
output = ["qc/{sample}.bin"]        # internal — not declared below
shell = "true"

# main.oxoflow
[[include]]
path = "qc.oxoflow"
namespace = "qc"
inputs  = ["raw/{sample}.fq.gz"]    # host MUST produce these
outputs = ["qc/{sample}_fastqc.html"]  # module EXPOSES these
params  = { threads = "4" }         # config defaults (host values win)

[[rules]]
name = "download"
output = ["raw/{sample}.fq.gz"]     # wires the declared input
shell = "true"
Field Type Description
inputs Array of String File patterns the host must wire into the module. A concrete declared input (no {-wildcard) that no rule produces fails validation with the gap named; wildcarded declared inputs are not statically checkable (a {sample} input forms no DAG edge — the placeholder only resolves at run time) — the host is responsible for wiring them
outputs Array of String Files the module exposes to the host. A declared output not produced by a module rule fails validation; a host rule reading a module-internal file that is not declared here produces an encapsulation warning. Wildcarded host inputs are matched structurally (identical literal prefix and suffix, e.g. qc/{sample}.tmp), so patterns that could address the same files warn too; patterns that resolve to different literals are not statically checkable
params Table Defaults for config keys the module reads, filled in profile-style (or_insert — explicit host values win). The module's own [config] entries fill any remaining gaps: a module declaring [config] trim_quality = "20" keeps that default when included without params, while any host value (its own [config] table or [[include]] params) still wins. A key referenced but defined nowhere remains an E005 error
name String Module identity for partial runs (run --module <name>). Defaults to the included file's stem (rules/20_germline.oxoflow → 20_germline), so existing composed workflows are addressable without changes
repo + ref String pair Pin the included path to a git repository at a tag/branch/commit: repo = "https://github.com/org/oxo-flow-modules", ref = "v0.23.2", path = "rules/qc.oxoflow". Any git URL works (https/ssh/file://); github.com clones fall back to China mirrors. Checkouts are cached under ~/.cache/oxo-flow/modules/<repo>@<ref>@<hash> (the hash disambiguates confusable repo/ref pairs; OXO_FLOW_MODULE_CACHE overrides the cache root) — versioned modules, reproducible composition

[workflow] — Metadata#

[workflow]
name = "my-pipeline"
version = "1.0.0"
description = "A short description"
author = "Your Name"
Field Type Required Default Description
name String Yes — Pipeline name
version String No "0.1.0" Semantic version
description String No — Human-readable description
author String No — Author name or email
interpreter_map Table No {} Custom interpreter mapping for script extensions
genome_build String No — Genome reference build identifier (e.g., "GRCh38", "hg38")
min_version String No — Minimum engine version required to load this workflow: an older engine refuses it at parse time (compares major.minor.patch; "0.21" equals "0.23.2")
format_version String No — Format specification version for compatibility
pairs_file String No — External TSV/CSV/JSON file defining experiment-control pairs
sample_groups_file String No — External TSV/CSV/JSON file defining sample groups
metadata_file String No — External TSV/CSV/JSON file of per-sample metadata columns, addressed as {meta.<column>} (see Sample Metadata)
pairs_pattern String No — File glob pattern for auto-discovering pairs (e.g., "aligned/{pair_id}/{exp}_vs_{ctrl}.bam")
sample_pattern String No — File glob pattern for auto-discovering samples (e.g., "raw/{sample}_R1.fastq.gz")
on_complete String No — Shell command run once after the whole workflow completes successfully (no failed rules). Rendered with {config.*} and the counters {succeeded}/{failed}/{skipped} plus {workdir}; runs in the workflow root with the workflow shell_prelude. Best-effort: a failing hook warns, never changes the run status.
on_error String No — Shell command run once after the workflow reaches a terminal state with at least one failed rule. Same rendering and best-effort semantics as on_complete.

Workflow-Level Terminal Hooks#

on_complete / on_error mirror the rule-level on_success / on_failure hooks at the whole-workflow level — the nf-core workflow.onComplete / snakemake onerror pattern (completion emails, cleanup, notifications):

[workflow]
name = "my-pipeline"
version = "1.0.0"
on_complete = "notify 'run done: {succeeded} succeeded, {failed} failed' --tag {config.project}"
on_error = "notify 'run FAILED: {failed} failed' --tag {config.project} && tar czf failed-run.tar.gz {workdir}/logs"

Both hooks run exactly once per oxo-flow run (the CLI) that reaches a terminal state (including --keep-going runs); on_complete fires only when no rule failed, on_error otherwise. The web UI executor has its own run loop and does not fire these hooks yet. Placeholders: every [config] key is available as {config.<key>}, plus {succeeded}, {failed}, {skipped}, and {workdir}. Unknown keys stay literal.

Custom Interpreters (interpreter_map)#

By default, oxo-flow auto-detects interpreters based on file extensions:

  • .py → python
  • .py3 → python3
  • .R, .r → Rscript
  • .sh, .bash → bash
  • .jl → julia
  • .pl → perl
  • .rb → ruby
  • .qmd, .Rmd, .rmd → quarto render
  • .ipynb → jupyter nbconvert --to notebook --execute
  • .smk → snakemake
  • .nextflow → nextflow run
  • .wdl → miniwdl run

You can override or extend this mapping in the [workflow] section:

[workflow]
name = "custom-interpreters"

[workflow.interpreter_map]
".m" = "octave"
".sas" = "sas"
".py" = "/opt/conda/bin/python"  # Override default

This mapping applies only to the script field.


[config] — Configuration Variables#

User-defined key-value pairs accessible in rules as {config.<key>}. Every config key can be overridden from the CLI.

Two forms are supported:

[config]
reference = "/data/ref/hg38.fa"             # plain value (implicit default)
samples_dir = "raw_data"
results_dir = "results"
min_quality = "30"

Values are TOML strings, integers, booleans, or arrays. String interpolation in rules uses {config.key} syntax; array values render as space-joined lists in the shell.

Several keys are engine-injected:

  • config.samples_list (all sample names, comma-joined) and config.pairs_list (all pair ids, comma-joined) — use them in expand_inputs to gather across samples or pairs without maintaining a parallel list by hand. A user-declared config.samples_list is union-merged with the discovered samples, and the two lists are also rewritten by --samples filtering.
  • config.samples_<group> (one per [[sample_groups]] entry).
  • config.<reference name> (each [[references]] entry's output) and, with reference_dir, the ten derived paths (reference_fasta, gene_annotation, bwa_index, bwamem2_index, bowtie2_index, star_index, hisat2_index, minimap2_index, gatk_dict, samtools_faidx). Explicit [config] entries with the same names are never overwritten.

Declarative Form (inline table)#

For user-facing parameters, use the declarative inline-table form. Declared keys become automatic CLI flags (--<key> <value>), can be required, and carry metadata:

[config]
database  = { required = true, help = "Path to the BLAST database" }
threshold = { default = "1e-5", help = "E-value cutoff" }
mode      = { default = "dna", choices = ["dna", "rna"], type = "string" }
samples   = { required = true, type = "path", must_exist = true }

Fields#

Field Type Description
default String Default value when not provided via CLI
required Bool If true, the value must be provided at runtime (default: false)
help String Human-readable description shown in error messages
sensitive Bool Mask the value as *** in captured rule output, recorded commands, checkpoint/report snapshots, and error summaries — before they reach the web UI or AI recovery (default: false). Every non-empty value is masked; values of 4+ characters also mask their base64, percent-encoded, and JSON-escaped forms
type String Expected value type for validation: "string", "int", "float", "bool", "path"
choices Array of String Allowed values (requires type = "string")
range String Numeric range "min..max" (requires type = "int" or "float")
must_exist Bool Path must exist on disk (requires type = "path")

Usage in Rules#

Values are accessible as {config.<name>} in shell commands, input/output paths, and when conditions:

[[rules]]
name = "blast"
shell = "blastn -db {config.database} -evalue {config.threshold} -query {input} -out {output}"

CLI#

# Provide a required value (direct flag form)
oxo-flow run pipeline.oxoflow --database refs/nt

# Attached form
oxo-flow run pipeline.oxoflow --database=refs/nt

# Key=value form directly after the workflow file
oxo-flow run pipeline.oxoflow database=refs/nt

# Override several values
oxo-flow run pipeline.oxoflow --database refs/nt --threshold 1e-3

# Legacy --arg form (still supported)
oxo-flow run pipeline.oxoflow --arg database=refs/nt --arg threshold=1e-3

Precedence#

CLI override > declared default > error if required and unset.

An override key that differs from a declared key only by hyphens vs underscores resolves onto the declared spelling: with threads_max = { default = "4" } in [config], --threads-max=8 sets threads_max (and the --threads-max 8 space form is accepted too). This matters most for sensitive keys — masking and the checkpoint snapshot are keyed on the declared spelling — but it applies to every declared key. Keys matching nothing declared are untouched.


[[references]] — Auto-Built Indexes & Reference Data#

[[references]] is for external, pre-staged reference artifacts: the engine builds them when missing, before the rule DAG runs. In-workflow derivations (e.g. a STAR index produced by an upstream rule) must flow through normal rule output/input edges instead — a reference entry cannot consume rule outputs because references are built before the DAG exists.

Declare reference artifacts (indexes, data files) that the engine auto-builds when missing. Each [[references]] entry specifies a source, output, and build command. The engine tracks built state in .oxo-flow/reference-checkpoint.json and never rebuilds unnecessarily: each entry stores a fingerprint of the definition plus the config values its build command references. The source file's content state joins the fingerprint: size and mtime for every source, plus a content hash for files up to 64 MiB — so for those files even a timestamp-preserving same-path replacement (cp -p, rsync -t, git checkout) triggers a rebuild; larger sources are guarded by size + mtime. Editing the build/source/output, changing a referenced config value, or touching the source file also triggers a rebuild, and rules that consume the artifact through declared input paths (plus their downstream) are invalidated so their outputs are regenerated.

When reference_dir is set, eight standard indexes are auto-derived without explicit [[references]] blocks: samtools faidx, BWA, BWA-MEM2, Bowtie2, minimap2, STAR, HISAT2, and the GATK sequence dictionary (see the auto-derivation table below).

Use --skip-ref-build to skip automatic reference building.

Syntax#

[[references]]
name = "bwa_index"
source = "{reference_dir}/genome.fa"
output = "{reference_dir}/bwa/genome.fa"
build = "mkdir -p {reference_dir}/bwa && bwa index -p {output} {source}"
threads = 8
memory = "8G"
description = "BWA index for alignment"

[references.environment]
conda = "envs/bwa.yaml"

Fields#

Field Type Required Description
name String Yes Unique name (used for checkpoint tracking)
source String No Source file for freshness checks
output String Yes Path to the built artifact
build String Yes Shell command to produce output from source
threads Integer No CPU threads for the build command
memory String No Memory limit (e.g., "64G")
description String No Human-readable description
environment Table No Environment providing the build tool — same spec as [rules.environment] (conda/mamba/docker/…). Without it the build runs in the bare system shell, so builds that need workflow tools (bowtie2-build, STAR --runMode genomeGenerate, …) must declare one. The environment is created on first use and shared with rules through the same cache.

Built-in Builder Templates#

Instead of a handwritten build shell, name one of the built-in templates — the engine expands it into a canonical command (any string containing whitespace or shell syntax is treated as a handwritten command and passes through unchanged):

Template Produces
samtools_faidx .fai index
picard_dict .dict sequence dictionary
bwa_index BWA's five index files (.amb/.ann/.bwt/.pac/.sa; the prefix is derived from output)
star_index STAR genome directory (output = --genomeDir)
bismark_index Bismark Bisulfite_Genome directory (source = FASTA directory)
[[references]]
name = "genome_bwa"
source = "refs/genome.fa"
output = "refs/genome.fa.bwt"
build = "bwa_index"
threads = 8

Unknown template names, missing outputs, and duplicate reference names are rejected at validate. Additionally, every reference injects a keyed config value — config.<name> = its output — so rules consume the artifact as {config.genome_bwa} without repeating paths.

Auto-Derivation from reference_dir#

When [config] contains reference_dir and no explicit [[references]] blocks are declared, the engine automatically derives eight standard references:

Name Output
samtools_faidx {reference_dir}/genome.fa.fai
bwa_index {reference_dir}/bwa/genome.fa.bwt
bwamem2_index {reference_dir}/bwamem2/genome.fa.0123
bowtie2_index {reference_dir}/bowtie2/genome.fa.1.bt2
minimap2_index {reference_dir}/genome.fa.mmi
star_index {reference_dir}/star/SAindex
hisat2_index {reference_dir}/hisat2/genome.fa.1.ht2
gatk_dict {reference_dir}/genome.dict

reference_dir also auto-derives config defaults such as reference_fasta ({reference_dir}/genome.fa) and gene_annotation ({reference_dir}/genes.gtf). Users can override any auto-derived reference by declaring it explicitly.

CLI#

# Auto-build missing references
oxo-flow run pipeline.oxoflow -j 16

# Skip reference building (assume pre-built)
oxo-flow run pipeline.oxoflow --skip-ref-build

[defaults] — Default Settings#

Applied to all rules unless explicitly overridden:

[defaults]
threads = 4
memory = "8G"
environment = { conda = "envs/base.yaml" }
shell_prelude = "set -euo pipefail"
Field Type Description
threads Integer Default CPU thread count
memory String Default memory allocation
environment Table Default environment specification
shell_prelude String Shell text prepended to every rule command, hook, reference build, and cluster job script on its own line (issue #92) — e.g. set -euo pipefail for fail-fast shells. Opt-in: empty by default, so existing workflows keep their exact command text. Applies inside environment wrappers (containers, conda) on every execution path. pipefail requires bash (the engine falls back to sh only when bash is absent)

Conda/mamba environment naming#

For file-backed env specs the conda/mamba env name is derived as <stem-or-yaml-name>-<hash8>, where hash8 is the first 8 hex chars of the SHA-256 of the YAML content (issue #159). Relative env file specs (conda, mamba, pixi, venv_requirements) resolve against the workflow directory — the single base for relative paths in the workflow file (issue #427); with --workdir the engine anchors them automatically, so rule shells running from the workdir still find the specs. venv (a directory the backend creates in the workdir) and conda_prefix/mamba_prefix (an environment store outside the workflow tree) stay relative to the workdir. Two workflows that ship different YAMLs under the same name therefore build into distinct envs instead of silently sharing one prefix; identical content keeps the same name, so specs still deduplicate. Non-file specs (plain names or inline strings) keep their plain name. Note: the first run after this convention shipped rebuilds envs once under the new names — remove the old plain-named envs with conda env remove -n <old> if disk space matters.


[report] — Report Configuration#

Configure report generation behavior. Reports are built by a pluggable section system — each section is produced by a ReportSectionGenerator. Use sections to control which generators run.

[report]
template = "report.html"
sections = ["universal", "workflow-info", "commands", "clinical-compliance"]
Field Type Description
template String Report template: the built-in name "report.html", or a template file path (workflow directory first, then cwd). Applies to HTML output only; a render failure warns and falls back to the default renderer
format Array Parsed but not supported yet — setting it makes report warn (or fail under --strict); select the output format with -f instead
sections Array Report sections to include. If empty (or omitted), all applicable generators run. Available built-in IDs: universal, execution-status, failure-diagnosis, clinical-compliance, workflow-info, commands, file-manifest, environment, metrics, sample-matrix, provenance, task-summary, software-versions, rule-captions, aggregate-metrics — 15 in total (run oxo-flow report --list-sections for the live list)

How Sections Work#

Each section ID maps to a registered ReportSectionGenerator:

Section ID Description When Active
universal Dashboard with QC indicators and task counts Always
workflow-info Name, version, author, config, sample/pair counts Always
commands Expanded shell commands for every rule Always
file-manifest Input and output file listings Always
environment Available backends and oxo-flow version Always
clinical-compliance ACMG/AMP classification, audit trail, biomarkers Always
execution-status Per-rule execution status and benchmark metrics Only with checkpoint
software-versions Declared software per rule, nf-core-style: docker image:tag, module tool/version, env files with sha256 hashes — static declaration, runtime package versions deliberately not recorded (see Reporting) When any rule declares an environment
rule-captions Per-rule report annotations rendered as markdown — see Rule Captions When any rule declares report
aggregate-metrics MultiQC-style sample × metric matrix across tools, plus *_mqc.json custom content When metric or custom-content files are found

The domain (DNA-seq, RNA-seq, epigenomics, or generic) is auto-detected from tool names in the workflow. Custom generators can be registered programmatically.

Rule Captions#

A rule can carry a report annotation — caption text rendered in the report's rule-captions section (the Snakemake report() equivalent). Two TOML forms:

[[rules]]
name = "qc"
shell = "fastp -i {sample}.fq"
# Inline caption (markdown allowed):
report = """
# QC summary
Reads trimmed with fastp; per-sample Q30 below.
"""

[[rules]]
name = "align"
shell = "bwa mem -t {threads} {input} > {output}"
# Structured form: a markdown/text file relative to the workdir,
# and/or an inline caption (at least one of the two):
report = { file = "notes/alignment.md", caption = "Alignment QC" }

At render time each executed rule instance contributes one markdown subsection, ordered by rule declaration order:

  • The checkpoint's execution-time caption wins — it is what the run actually read (report.file content is captured when the rule completes, so a later edit to the file does not silently change a finished report).
  • Without an execution record (dry run, or a checkpoint from before this feature), the declared annotation is resolved at report time: inline caption first, then report.file relative to the workdir.
  • A missing or unreadable report.file never fails the report — the rule simply contributes no caption.

Captions are text/markdown only; embedding figures is not supported in v1.


[[rules]] — Rule Definitions#

Each [[rules]] entry defines a pipeline step. The double brackets indicate a TOML array of tables.

Basic example#

[[rules]]
name = "align"
input = ["{sample}_R1.fastq.gz", "{sample}_R2.fastq.gz"]
output = ["aligned/{sample}.bam"]
environment = { conda = "envs/alignment.yaml" }
shell = "bwa mem -t {threads} {config.reference} {input} | samtools sort -o {output}"

[rules.resources]
threads = 16
memory = "32G"

All fields#

Field Type Required Description
name String Yes Unique rule identifier
input Array of strings No Input file paths. Optional (defaults to empty) — a rule with no input is legal (e.g. a pure producer or a setup step; pair it with depends_on when ordering matters)
output Array of strings No Output file paths. Optional (defaults to empty) — a shell-only rule with no declared outputs validates and runs; pair with depends_on for ordering
output_pattern String No Runtime-discovered output pattern: outputs enumerated by a filesystem scan after the rule completes; downstream consumers of its wildcard instantiate per discovered value. A [[values]] wildcard in the pattern is a plan-time fan-out trigger instead (issue #296). Mutually exclusive with output/transform/input_groups — see Runtime-Discovered Outputs
shell String No Shell command to execute
script String No Script file path (auto-detects interpreter)
description String No Human-readable description of what this rule does
report String or Table No Report annotation rendered by the rule-captions report section. Inline text: report = "# QC summary\nReads trimmed with fastp." (markdown allowed). Or a table: report = { file = "notes/qc.md", caption = "Alignment QC" } where file is a markdown/text file resolved relative to the workdir (at least one of file/caption required) — see Rule Captions
threads Integer No (Deprecated) CPU threads — use resources.threads instead. Still honored, but oxo-flow lint flags it with a W025 suggestion to move it under [rules.resources]
memory String No (Deprecated) Memory allocation — use resources.memory instead. Still honored, but oxo-flow lint flags it with a W025 suggestion to move it under [rules.resources]
resources Table No Full resource specification (threads, memory, gpu, disk, time_limit, partition, groups)

Run-path field validation (#523). Rule names (charset: alphanumeric, _, -, :), resources.memory/memory formats (<number><G|M|K|T>), thread counts and retry_delay formats are validated when the workflow is parsed — an invalid value fails the run before any expansion or rendering, so values from a third-party [[include]] can never reach the rendered docker/scheduler/shell command line. Design-level checks (GPU/backend pairing, output_pattern consistency) stay at expansion time where their context-aware diagnostics run. | environment | Table | No | Environment specification | | transform | Table | No | Unified scatter-gather operator (split → map → combine) | | when | String | No | Conditional expression — skip rule when false. May call read-only runtime functions (reads_count, wc_lines, file_size, file_exists, regex_extract) that inspect files in the run workdir; see Data-Dependent when Gates | | envvars | Table | No | Dictionary of environment variables to inject | | params | Table | No | User-defined parameters for shell templates | | pre_exec | String | No | Command to run before the main shell command | | on_success | String | No | Command to run after rule succeeds | | on_failure | String | No | Command to run after rule fails (all retries exhausted) | | retries | Integer | No | Number of retry attempts on failure (default: 0) | | interpreter | String | No | Explicit interpreter for script execution | | checkpoint | Boolean | No | Enable checkpoint re-entry after this rule completes (requires checkpoint_manifest) | | checkpoint_manifest | String | No | Path (relative to workdir, {config.x}-expanded) of the TOML manifest the checkpoint rule writes at runtime to declare new samples | | scatter | Table | No | Fan-out parallel execution over a variable with optional gather | | expand_inputs | Table | No | Cartesian product expansion of input patterns | | input_groups | Array of tables | No | Group files by a pattern wildcard into one per-group instance (the Nextflow groupTuple pattern) — see Input Groups | | priority | Integer | No | Execution priority (higher = runs first; default: 0) | | target | Boolean | No | Default target: run/dry-run without an explicit -t/--module builds the rules marked target = true and their transitive producers; with no marked rule the full DAG runs (an explicit -t always wins) | | required | Boolean | No | Pipeline fails if this rule fails, even without downstream deps | | optional | Boolean | No | Rule is skipped (no error) when its inputs don't exist — literal globs count as missing only when they match nothing | | benchmark | String | No | Per-instance benchmark file: after the rule completes, the engine writes its sampled metrics (wall time, peak RSS, memory limit, CPU seconds, retries) as JSON to this path — engine wildcards in the path resolve per instance, and a failed write never fails the rule | | log | String | No | Log file path for rule execution output. The engine creates its parent directory and expands {config.x}/wildcards in the path, but populating the file is the rule's job — typically via 2> {log} in the shell, so it carries content only when the tool writes to stderr. A tool that self-logs to its own location leaves the declared log empty; that never means diagnostics were lost — the captured stdout/stderr tails are stored in the checkpoint's rule_runs regardless, and a succeeded rule whose declared log and both captured tails are all empty earns one Info hint that the tool likely self-logs (issue #832) | | group | String | No | Reserved metadata: cluster array grouping keys on the rule's template name by design; this label is round-tripped for external tooling | | cache_key | String | No | Content-addressed output reuse: cached outputs are restored when the key, inputs, outputs, and rendered command hash identically to a previous run (issue #194 §2.3) | | input_function | String | No | Parsed but not yet called — no dynamic input resolution is performed today | | rule_metadata | Table | No | Domain-specific metadata (assay, organism, etc.) — a round-tripped surface for external tooling; the engine itself does not render it | | env_group | String | No | Reference to a named environment in [env_groups] | | depends_on | Array | No | Explicit rule-level dependencies (by rule name) | | extends | String | No | Inherit settings from a base rule | | retry_delay | String | No | Delay between retries (e.g., "5s", "30s", "2m") | | temp_output | Array | No | Scratch artifacts the rule itself overwrites (e.g. .tmp.bam). Cleaned when the rule fails (so a stale partial never masquerades as a completed run); they persist on success — mark temporary = true for success-path deletion | | temporary | Boolean | No | Delete the rule's outputs after a fully successful run once every dependent has completed, recording a tombstone so a future run regenerates them on demand (leaf rules keep their outputs) | | scratch | Boolean | No | Execute in an isolated per-instance scratch directory: inputs render as absolute paths, declared outputs written there move back to the main workdir and are verified, and the scratch is removed on success (preserved with a path note on failure). Use for tools that write fixed filenames or pollute the workdir | | protected_output | Array | No | Outputs the engine must never destroy: failure invalidation skips them (no delete, no move-aside to <name>.oxo-failed) and oxo-flow clean refuses them with a (protected — skipped) diagnostic, even with --force. Three pattern forms are honored: exact paths, {wildcard} patterns (results/{sample}.bam), and shell globs (results/*.bam — * does not cross /). Advisory in the remaining sense that the rule's own command can still overwrite them on a rerun | | tags | Array | No | Categorization tags (e.g., ["qc", "alignment"]) |

| ancient | Array | No | Inputs excluded from freshness tracking and input manifests — they never trigger re-execution (reference files; an engine wildcard in a declaration matches the expanded input structurally) | | localrule | Boolean | No | Parsed but not enforced yet — mixed local+cluster execution needs a scheduler-side local executor and live cluster verification; every schedulable rule can currently be submitted. Do not rely on it (see Resource Tuning) | | format_hint | Array | No | Parsed and validated but not yet used by the engine (planned for I/O optimization) | | pipe | Boolean | No | Parsed and validated but not yet used by the engine (planned for FIFO streaming; no streaming is performed today) | | checksum | String | No | Parsed but not yet enforced — provenance checksums (sha256) are computed for all outputs regardless; no per-rule verification happens today | | resource_hint | Table | No | Resource estimation hints for dynamic scheduling |

Note: When a rule declares outputs, at least one of shell, script, or transform must be provided. If both shell and script are defined, they execute sequentially: shell first, then script.

Environment specification#

Exactly one backend per rule (uncomment the one you need):

[[rules]]
name = "example"
# Conda
environment = { conda = "envs/tools.yaml" }

# # Pixi
# environment = { pixi = "envs/pixi.toml" }

# # Docker
# environment = { docker = "biocontainers/bwa:0.7.17" }

# # Singularity
# environment = { singularity = "docker://biocontainers/bwa:0.7.17" }

# # Python venv
# environment = { venv = ".venv/" }

# # HPC modules
# environment = { modules = ["gcc/11.2.0", "openmpi/4.1.1"] }

# # Mamba / micromamba (auto-detects binary, same YAML format as conda)
# environment = { mamba = "envs/qc.yaml", mamba_prefix = ".oxo-flow/envs" }

# # Conda with custom prefix
# environment = { conda = "envs/qc.yaml", conda_prefix = ".oxo-flow/envs" }

# # venv with custom requirements file
# environment = { venv = ".venv/", venv_requirements = "envs/dev-requirements.txt" }

# # Reference a named environment group (defined in [env_groups])
# env_group = "qc_env"

Container shell: docker/singularity rules run under bash inside the container when the image provides it (falling back to sh otherwise). nf-core-derived images ship bash, and their scripts rely on bash features (set -o pipefail, [[ ]]) that the container default sh (often dash) rejects — the re-exec shim keeps these scripts working unmodified.

Inline conda package lists: besides a YAML path, a conda/mamba spec may be a channel-qualified package list — the concise form for rules whose environment is a handful of pinned tools:

[[rules]]
name = "qc"
environment = { conda = "bioconda::fastp=0.23.4 bioconda::samtools=1.24" }

Whitespace- or comma-joined packages are the same list. Every package must be channel-qualified (channel::name=version) and match the strict package charset — the tokens are rendered as separate conda create arguments, so shell metacharacters are still rejected (E016). The env name is derived from the first package plus a content hash of the whole list: identical package sets share one environment, a changed set builds a fresh one. Specs without :: (paths, lockfiles) keep the YAML-path semantics.

Singularity spec shapes: a singularity spec is either a pull URI (docker://, library://, oras://, https://) — the engine pulls and converts it to a SIF in the workdir on first use, reusing the SIF when it exists — or a local SIF path (e.g. a site image store's /data/images/bwa_0.7.17.sif), which is used as-is with no pull step. A local path that does not exist fails the rule with a clear diagnostic.

GPU passthrough (gpus): environment = { docker = "nvcr.io/nvidia/clara-parabricks:4.4.0-1", gpus = "all" } passes the host's GPU devices to the container via docker's --gpus <value> (also "0" or "device=0,1"). Requires nvidia-container-toolkit on the host. Declaring gpus without a docker image is a validation error — the flag would otherwise be silently ignored on system/conda backends. (Singularity --nv passthrough is not implemented yet; gpus with a singularity image is rejected for now.)

Named Environment Groups ([env_groups])#

Instead of repeating the same environment spec across multiple rules, define named groups once in [env_groups] and reference them via env_group:

[env_groups.qc_env]
conda = "envs/qc.yaml"

[env_groups.align_env]
conda = "envs/alignment.yaml"

[[rules]]
name = "fastqc"
env_group = "qc_env"
input = ["raw/{sample}.fastq.gz"]
output = ["qc/{sample}_fastqc.html"]
shell = "fastqc {input} -o qc/"

Rules using env_group inherit the full environment specification from the named group. The env_group spec takes precedence over an inline [rules.environment]; the inline spec is only used when no env_group is set (or the named group does not exist), followed by [defaults] environment.

Environment Variables (envvars)#

Inject rule-specific environment variables directly into the execution context:

[[rules]]
name = "deep_learning"
shell = "python train.py"

[rules.envvars]
CUDA_VISIBLE_DEVICES = "0"
PYTHONPATH = "./src"

Variables defined here are available to the main shell command as well as all lifecycle hooks (pre_exec, etc.).

Parameters (params)#

Define custom variables for use in shell templates. Unlike [config], which is global, params are specific to a single rule and take precedence during interpolation:

[[rules]]
name = "count_reads"
shell = "samtools view -c -q {params.min_qual} {input} > {output}"

[rules.params]
min_qual = 20

Script Execution (script)#

The script field allows you to execute external script files (Python, R, etc.) with automatic interpreter detection.

[[rules]]
name = "analyze"
script = "scripts/analysis.py --min-quality {params.q}"
interpreter = "python3" # Optional: overrides auto-detection

Interpreter Detection Order:

  1. Explicit interpreter field on the rule.
  2. Custom [workflow.interpreter_map] in the metadata.
  3. Built-in defaults based on file extension.

Lifecycle Hooks#

Hooks allow you to run auxiliary logic at different stages of a rule's life:

[[rules]]
name = "process_data"
shell = "python process.py"
pre_exec = "mkdir -p tmp_workspace"
on_success = "echo 'Success!' | slack-notify"
on_failure = "rm -rf tmp_workspace && echo 'Cleanup done'"
retries = 3
  • pre_exec: Runs before the main command. If it fails, the rule is aborted.
  • on_success: Runs only after the main command completes with exit code 0.
  • on_failure: Runs if the main command fails, after all retries have been exhausted.

Hook commands support the same {config.x} / {input} / {output} placeholder expansion as shell.

Output validation: After a rule's shell command exits successfully (exit code 0), oxo-flow verifies that all declared output files exist on disk. If any declared output is missing — for example, because a multi-step shell script's cleanup step masked an earlier tool failure — the rule is treated as failed (with exit code sentinel -1) and the on_failure hook runs instead of on_success. This prevents silently broken rules from being checkpointed as "completed" and skipped on resume.


Resources (extended)#

For rules needing GPU, disk, or time limits, use the resources sub-table:

[[rules]]
name = "gpu_task"
input = ["data.h5"]
output = ["model.pt"]
shell = "python train.py"

[rules.resources]
threads = 8
memory = "64G"
gpu = 1
disk = "200G"
time_limit = "48h"

Declared disk requirements are checked against the working directory's free space before a run starts — a shortfall prints a warning (the run proceeds; the warning exists so long jobs don't fail mid-pipeline).

Field Type Example Description
threads Integer 8 Number of CPU threads
memory String "16G" Memory allocation
gpu Integer 1 Number of GPUs
disk String "200G" Local disk space
time_limit String "48h" Wall-time limit
partition String "gpu" HPC partition/queue to submit to
groups Table {db_conn = 1} Resource group consumption tracking

Resource Management#

Declaration vs Enforcement#

oxo-flow tracks declared resources for scheduling but does not strictly enforce them in local execution. On HPC clusters, resources are enforced by the scheduler.

Local execution:

  • Resources are tracked to prevent over-allocation
  • Warnings emitted when declaring resources exceeding system capacity
  • Jobs may oversubscribe if user intentionally requests more than available

HPC clusters:

  • Resources translated to scheduler directives (SLURM, PBS, SGE, LSF)
  • Scheduler enforces limits - jobs requesting more than allocated will fail

Platform Detection#

Platform Thread Detection Memory Detection
Linux num_cpus crate sysinfo crate
macOS num_cpus crate sysinfo crate

Validation Warnings#

When a rule declares resources exceeding system capacity, oxo-flow emits warnings during validation but does not block execution:

⚠️  rule 'bwa_align' requests 128 threads but system has 64 (will oversubscribe)
⚠️  rule 'big_sort' requests 128GB but system has 32GB (may OOM)

This allows intentional oversubscription for testing or when user knows better.

Cleanup Behavior#

oxo-flow automatically cleans up temporary outputs:

| Scenario | Cleanup | |---|---|---| | Success + temp_output | Kept — no success-path cleanup runs; mark temporary = true instead if you want outputs deleted after the run | | Failure + temp_output | Cleaned to prevent stale partial files | | Failure + declared output | Partial outputs created by the failed attempt are deleted; all pre-existing declared outputs are moved aside as <name>.oxo-failed (modified or untouched — a stale file from an earlier run era would otherwise block the re-execution the recorded failure schedules with File exists errors; issue #756) | | Transform with cleanup=true | Chunk files cleaned after the whole run finishes successfully (kept on failed runs for debugging; re-runs recompute the map rules) | | Success + temporary = true | Outputs deleted after the run once every dependent rule has completed; a tombstone is recorded in the checkpoint so a later run that needs the outputs regenerates the rule first (lazy cascade-up). Leaf rules (no dependents) keep their outputs. For output_pattern rules, deletion is per recorded instance (only the concrete paths that ran and completed — issue #844); patterns that cannot be expanded produce a warning naming the rule instead of a silent no-op. If a dependent has not completed yet, the outputs are kept and a warning names the rule and its pending dependents — a later run (or resume) cleans them up once the dependents finish |

temporary is for whole intermediates (e.g. multi-GB per-sample BAMs kept only until the queue-level callers finish): the deletion is checkpoint-aware, so a plain re-run skips the rule and does NOT regenerate the file — it comes back only when a dependent actually needs it again.

Failed-rule output invalidation. A rule that fails mid-write must never leave partial outputs that the freshness gate would mistake for a completed run. Before executing, the engine snapshots each declared output's existence and modification time; when the rule fails (non-zero exit, timeout, or missing outputs after an exit-0 shell), the declared outputs are invalidated: files created during the attempt are deleted, and every pre-existing declared output moves aside as <name>.oxo-failed — modified or untouched (issue #756). The recorded failure schedules a re-execution, and a stale file surviving at the declared path would block that re-run inside the script (ln: File exists, mv: cannot move ... File exists); move-aside keeps it recoverable, and protected_output remains the escape hatch for files that must survive any failure. The next run then re-executes the rule instead of skipping it as up-to-date. Note this covers failures the engine observes; a hard kill (SIGKILL) mid-rule can still leave outputs behind, since no cleanup code runs — the --rerun flag is the escape hatch for that case.

Timeout Enforcement#

On Unix systems (Linux, macOS), a timeout kill walks the rule's process subtree rather than a process group — rules deliberately run inside the run's process group ("one run = one process group"), so the engine snapshots parent→child links via sysinfo and signals each descendant individually, deepest first: SIGTERM to the whole subtree, a 10-second grace poll for survivors, a re-scan for descendants spawned during the window, then SIGKILL to whatever remains (issue #194):

[rules.resources]
time_limit = "4h"  # SIGTERM → grace → SIGKILL escalation after 4 hours

GPU Specification#

For detailed GPU requirements:

[rules.resources.gpu_spec]
count = 2
model = "A100"           # SLURM: --gres=gpu:a100:2
memory_gb = 40           # SLURM: --mem-per-gpu=40G
compute_capability = "8.0"  # For filtering (not scheduler directive)

Note: PBS/SGE GPU syntax varies by site. Use extra_args for site-specific flags.

Resource Hints#

When exact requirements unknown, provide hints for estimation:

[rules.resource_hint]
input_size = "medium"     # small (~1GB), medium (~10GB), large (~100GB), xlarge (~500GB)
memory_scale = 2.0        # Estimated memory = input_size × scale
runtime = "slow"          # reserved: parsed, not yet consumed by the scheduler
io_bound = true           # reserved: parsed, not yet consumed by the scheduler

Memory estimation formula: estimated_mb = input_size_mb × memory_scale. The runtime and io_bound hints are reserved scheduler knobs: only input_size and memory_scale are consumed today.


Script Execution#

shell commands run through bash (falling back to sh when bash is unavailable), so bashisms — brace expansion, [[ ]] conditionals, process substitution, $(( )) arithmetic — work in rule shells.

Script Field#

Execute a script file instead of (or in addition to) a shell command:

[[rules]]
name = "analysis"
input = ["data.csv"]
output = ["results.json"]
script = "scripts/analyze.py"  # Auto-detects interpreter from extension

When both shell and script are defined, they execute sequentially: shell first, then script.

[[rules]]
name = "qc_and_report"
shell = "fastqc {input} -o qc/"
script = "reports/qc_report.qmd"  # Runs after shell completes

Interpreter Detection#

oxo-flow automatically detects the interpreter from script file extension:

Extension Interpreter Notes
.py python Python script
.R / .r Rscript R script
.jl julia Julia script
.sh / .bash bash Shell script
.pl perl Perl script
.rb ruby Ruby script
.qmd quarto render Quarto document
.Rmd / .rmd quarto render R Markdown
.ipynb jupyter nbconvert --to notebook --execute Jupyter notebook
.smk snakemake Snakemake workflow
.nextflow nextflow run Nextflow script
.wdl miniwdl run WDL workflow

Explicit Interpreter Override#

Override auto-detection with interpreter field:

[[rules]]
name = "custom_python"
script = "analyze.py3"
interpreter = "python3.11"  # Override default python

Custom Interpreter Map#

Configure custom interpreter mappings at workflow level:

[workflow]
name = "pipeline"

[workflow.interpreter_map]
".m" = "octave"        # MATLAB/Octave
".sas" = "sas"         # SAS
".do" = "stata-mp"     # Stata
".stan" = "cmdstan"    # Stan

Additional Rule Fields#

Output Management#

Field Type Description
temp_output Array Scratch artifacts cleaned when the rule fails; persist on success (see Cleanup Behavior)
protected_output Array Outputs the engine never deletes or moves aside on failure; clean also refuses them. Exact paths, {wildcard} patterns, and shell globs are all honored (see the rule-field table above)
temporary Boolean Delete the rule's outputs after a fully successful run once every dependent has completed (tombstone + lazy regeneration; leaf rules keep outputs)
[[rules]]
name = "align"
output = ["aligned/{sample}.bam", "aligned/{sample}.bam.bai"]
temp_output = ["aligned/{sample}.tmp.bam"]  # Cleaned if the rule fails; kept on success
temporary = true                             # Delete aligned/*.bam once all callers finish

Directory outputs#

Directories are valid outputs — declare them with a trailing slash (or a bare directory name). Everything a rule writes inside that directory counts as produced, which is the idiomatic way to capture tool-generated side directories such as multiqc_data/ and multiqc_plots/:

[[rules]]
name = "multiqc"
output = [
    "results/multiqc/multiqc_report.html",
    "results/multiqc/multiqc_data/",
    "results/multiqc/multiqc_plots/",
]

Declaring directories matters for two reasons: downstream rules and the report see the whole directory as available, and the executor checks declared outputs after the shell exits — a rule that exits 0 while a declared output (file or directory) is missing is marked as failed. Conversely, do not declare directories the shell only uses as scratch: an undeclared directory that the rule recreates on re-run is lost on resume.

Execution Control#

Field Type Default Description
depends_on Array — Explicit rule dependencies (not inferred from files)
localrule Boolean false Parsed but not enforced yet — every schedulable rule can currently be submitted to the cluster

| checkpoint | Boolean | false | Enable checkpoint re-entry (requires checkpoint_manifest; the rule must not use {sample}/{group}) |

DAG edge inference — edges form when a rule's input paths match another rule's output paths, for every input form: template paths, expand_inputs patterns (both the expanded concrete paths at run time and the raw patterns in the template-level graph — multiqc-style aggregators therefore connect to their contributors even in graph -f dot without --expanded), glob patterns (raw/*.fastq.gz), and directory inputs (every producer writing under that directory). Rules that declare the SAME output string (shared bins/staging directories — e.g. two assembler variants emitting into one GenomeBinning/<tool>/bins dir) link ALL of their producers: every consumer of the shared path is ordered after every writer. depends_on remains available for ordering that files cannot express, and duplicate edges are deduplicated.

[[rules]]
name = "setup"
shell = "mkdir -p results"
depends_on = []  # Run first, before file-based dependencies

[[rules]]
name = "local_only"
shell = "echo 'local task'"
localrule = true  # NOTE: not yet enforced — every schedulable rule can be submitted

Input Groups (input_groups)#

Group many files of one sample (or any pattern key) into a single per-key rule instance whose {input} renders all of them — the Nextflow groupTuple pattern. Unblocks lane-merging (cat a sample's lane files in one step), library merging, and multi-file per-sample catalog steps.

[[rules]]
name = "lanemerge"
input_groups = [
    { pattern = "results/adapterremoval/{sample}_{lane}_R1.fastq.gz",
      group_by = "sample",
      keep = "lane" }
]
output = ["results/merged/{sample}_R1.fastq.gz"]
shell = "cat {input} > {output}"
  • pattern — path glob with {wildcard} placeholders, relative to the workflow file's directory (the same root sample_pattern scans). A bare * outside {wildcard} spans is a segment-local glob (issue #246): raw/{sample}_*_R1.fastq.gz groups every lane file of a sample without enumerating lane names — the CAT_FASTQ shape, where the per-sample file count is not known ahead of time and no [[values]] table can list it. ** (cross-segment recursion) is rejected. Note a {wildcard} span may span a separator by design (\S+); the segment locality applies to bare * only.
  • group_by — the wildcard that names each group. One instance is created per discovered group value. May also be a metadata column, group_by = "meta.<column>" — see below.
  • keep — optional restriction: which other pattern wildcards are bound to the group's first (sorted) file and become available as {wildcard} placeholders in output/shell. Omitted = all other wildcards. A string or an array of strings is accepted. Not allowed with group_by = "meta.<column>" (metadata-grouped instances bind only the column name).
  • {input} — all matched files of the group, space-joined in stable sorted order (files sort by path; the first sorted file supplies the keep bindings, so results never depend on readdir order).
  • {input_group.<wildcard>} — a captured per-group value as a list (space-joined), e.g. {input_group.lane} renders L1 L2. These are lists: use them in shell loops/commands, not as single paths.
  • Files are discovered at plan time from two sources: files already on disk under the workflow root, and the literal outputs of upstream rules — so a producer rule declared earlier in the file (e.g. a per-{sample}_{lane} rule) flows straight into the group's {input}, with DAG edges inferred from those plan-time-known paths.
  • Regular input entries coexist: group files render first, declared inputs append after them in {input}.
  • A group key with zero files instantiates nothing — the rule is skipped for that key with a warning (no instance = nothing to run).
  • When the workflow declares a sample domain ([[sample_groups]], sample_pattern, or [[pairs]]), group_by = "sample" group keys must be in that domain — keys outside it are pruned with a warning (#246). A key whose value is an underscore-suffixed form of a declared sample (S1_REP1 for declared S1) is kept: replicate-suffix naming folds by base id, the same convention the group pattern itself uses.

v1 constraints: at most one input_groups entry per rule; ** recursion is not supported in pattern (bare * globs within one segment are supported); producers must be declared before their input_groups consumer; the keep wildcard bindings come from the first sorted file only (the captured per-group values of every file stay available via {input_group.*}); {config.x} placeholders in pattern resolve against the workflow config. A producer whose output is a directory pattern (e.g. results/star/{sample}/) contributes only the directory path itself to the candidate pool — files inside it do not match input_groups patterns (declare the producer's files as outputs to group them).

# A multiqc-style aggregator over grouped files: one instance per
# sample, receiving every fastQC zip of that sample.
[[rules]]
name = "multiqc"
input_groups = [
    { pattern = "qc/fastqc_{sample}_{replicate}.zip", group_by = "sample" }
]
output = ["qc/multiqc_{sample}.html"]
shell = "multiqc {input} -o qc -n {output}"

Grouping by metadata column#

group_by = "meta.<column>" groups the matched files by a metadata_file column value instead of a pattern wildcard — the chipseq multi-antibody pattern, where the BAMs of several samples carry the same antibody and one peak-calling instance per antibody is wanted:

[workflow]
metadata_file = "samples.tsv"   # sample  antibody
#                                S1       H3K27ac
#                                S2       H3K27ac
#                                S3       Input

[[rules]]
name = "peaks"
input_groups = [
    { pattern = "results/mapping/{sample}.bam", group_by = "meta.antibody" }
]
output = ["peaks/{antibody}.bed"]
shell = "macs2 callpeak -t {input} -n {antibody} --outdir peaks && echo {input_group.sample}"
  • One instance per distinct column value: peaks_H3K27ac receives results/mapping/S1.bam results/mapping/S2.bam and peaks_Input receives results/mapping/S3.bam.
  • The group key binds under the column name ({antibody}), not {sample} — a metadata-grouped instance has no single sample. Files whose row has an empty column value are skipped.
  • Outputs must not reference pattern wildcards ({sample} in an output is a plan-time error); the per-group sample names stay available as {input_group.sample}.
  • keep is rejected with group_by = "meta.<column>" — the column name is the instance's only binding.

Retry Configuration#

Field Type Default Description
retries Integer 0 Number of automatic retry attempts
retry_delay String — Delay between retries ("5s", "30s", "2m")
[[rules]]
name = "network_task"
shell = "curl https://api.example.com/data"
retries = 3
retry_delay = "30s"

Input/Output Hints#

Field Type Description
ancient Array Inputs excluded from freshness tracking and input manifests — they never trigger re-execution
format_hint Array Parsed but not yet used by the engine (planned for I/O optimization)
pipe Boolean Parsed but not yet used by the engine (planned for FIFO streaming; no streaming is performed today)
checksum String Parsed but not yet enforced — provenance checksums (sha256) are computed for all outputs regardless
[[rules]]
name = "align"
input = ["reads/{sample}.fastq.gz", "ref/hg38.fa"]
ancient = ["ref/hg38.fa"]  # never triggers rebuild, however new the file is
# format_hint and checksum are accepted but currently have no effect

Organization#

Field Type Description
tags Array Categorization tags (["qc", "alignment"])
extends String Base rule to inherit settings from
[[rules]]
name = "align_default"
tags = ["alignment", "production"]

[rules.resources]
threads = 8
memory = "32G"

[[rules]]
name = "align_fast"
extends = "align_default"  # Inherits threads, memory, tags

[rules.resources]
threads = 16  # Override inherited value

Priority and Targeting#

Field Type Description
priority Integer Execution priority (higher runs first; default: 0)
target Boolean Default target: built by run/dry-run when no explicit -t/--module is given (marked rules + their transitive producers; none marked → full DAG)
[[rules]]
name = "critical_step"
priority = 10   # Runs ahead of lower-priority rules
target = true   # built by `run`/`dry-run` when no -t is given

Optional and Required Rules#

Field Type Description
optional Boolean or "any" true: the rule is skipped (no error) when ANY declared input is missing — literal globs count as missing only when they match nothing. "any": the rule runs when AT LEAST ONE declared input exists and is skipped only when none do — the alternative-input pattern, e.g. a consumer that accepts whichever mapper/aligner wrote the file. false (default): every input must exist
required Boolean If false, a failure of this rule does not fail the run: the failure is recorded and listed, its dependents are blocked, and the run exits 0 (best-effort / continue-on-error, e.g. QC steps). Defaults to true — required failures fail the run (or, with --keep-going, keep it running but still exit non-zero with the failures listed: keep-going changes scheduling, never the verdict)
[[rules]]
name = "experimental"
optional = true     # Skip if input data is absent
required = false    # Failure is reported but does not fail the run

[[rules]]
name = "filter_bam"
# Run when any of the mapper-specific BAMs exists (bwa writes *_PE.mapped.bam,
# bwamem/bt2/circularmapper write {sample}.mapped.bam); skip only when none do.
optional = "any"
input = [
    "results/mapping/bwa/{sample}_PE.mapped.bam",
    "results/mapping/bwa/{sample}.mapped.bam",
    "results/mapping/bt2/{sample}.mapped.bam",
]

When a rule is skipped because none of its optional inputs exist, non-optional dependents are skipped too ("blocked by skipped upstream dependency") instead of executing a doomed shell command on missing files.

Logging and Benchmarking#

Field Type Description
log String File path for capturing rule stdout/stderr
benchmark String Per-instance benchmark file — the engine writes its sampled metrics as JSON here after the rule completes
[[rules]]
name = "align"
log = "logs/align_{sample}.log"
benchmark = "benchmarks/align_{sample}.json"  # per-instance JSON, written by the engine

Job Grouping and Caching#

Field Type Description
group String Reserved metadata: cluster array grouping keys on the rule's template name by design; this label is round-tripped for external tooling
cache_key String Content-addressed output reuse: cached outputs are restored when the key, inputs, outputs, and rendered command hash identically to a previous run (issue #194 §2.3)
[[rules]]
name = "variant_call"
group = "variant_calling"       # reserved metadata (grouping keys on the template name)
cache_key = "vc_v2.0"           # Content cache key — outputs are reused when content is identical

Dynamic Input Resolution#

Field Type Description
input_function String Parsed and included in the rule fingerprint but not yet called — no dynamic input resolution is performed today (use expand_inputs or output_pattern for dynamic inputs)

Arbitrary Metadata#

Field Type Description
rule_metadata Table Domain-specific metadata (assay type, organism, protocol, etc.) — a round-tripped surface for external tooling; the engine itself does not render it
[[rules]]
name = "wgs_align"
[rules.rule_metadata]
assay = "WGS"
organism = "Homo sapiens"
protocol = "Illumina_NovaSeq_6000"

Scatter-Gather (Legacy)#

The scatter field provides fan-out parallelism over a variable with optional gather. For new workflows, prefer the unified transform operator.

Field Type Description
scatter.variable String Variable to scatter over (e.g., "chr")
scatter.values Array Values to scatter across
scatter.values_from String Config variable reference for values
scatter.gather String Name of the gather rule
[[rules]]
name = "per_chr"
scatter = { variable = "chr", values = ["chr1", "chr2", "chr3"] }

A rule whose scatter fans out over zero values — declared values = [], or values_from referencing a config key that resolves to an empty list — produces zero instances and can never run. oxo-flow validate therefore skips the input-existence check (E010/W020) for such rules and excludes their output templates from the W033 output-collision lint, mirroring the when-gated-off exemption: a rule that cannot run must not fail validation over inputs it will never read. An unresolvable values_from (missing config key, non-string/array value) is deliberately NOT treated as empty — the executor expands zero instances silently in that case, so validate keeps the input-existence check to surface the typo as a missing-input diagnostic instead of silently passing (issue #616; the stricter [[values]] tables and transform.split reject an unresolvable reference outright).

Expand Inputs#

The expand_inputs field generates additional input combinations via Cartesian product expansion.

Field Type Description
expand_inputs[].pattern String Input pattern with variables
expand_inputs[].variables Table Variable name → config reference (TOML array or comma-separated string)
[[rules]]
name = "multi_ref_align"
expand_inputs = [
  { pattern = "refs/{ref_genome}.fa", variables = { ref_genome = "config.ref_genomes" } }
]

Variables may also reference the [config] section — either a TOML array or a string:

[[sample_groups]]
name = "cohort"
samples = ["NA12878", "NA12879", "NA12880"]

[[rules]]
name = "combine_gvcfs"
input = []
expand_inputs = [
  { pattern = "variants/{sample}.g.vcf.gz", variables = { sample = "config.samples_list" } }
]

Config reference semantics:

Config value Expands to
["a", "b"] (array) a, b
"a" (string, no comma) a
"a,b,c" (comma-joined string) a, b, c (split on commas, trimmed)
["a,b"] (single-element array) a,b as one value — escape hatch for comma-containing strings

Inline literals (a plain string written directly in variables, e.g. sample = "S1,S2") are always treated as one value and are never split — the comma-splitting applies only to config.* references.

Comma-joined strings are how the engine injects merged sample lists (config.samples_list, config.samples_<group_name> — see Merging Multiple Sample Sources), so they can be referenced directly here.

The aggregation idiom — multi-sample merge/collapse steps consume the expanded inputs through {input}, which renders space-joined:

[[rules]]
name = "merge_bams"
expand_inputs = [
  { pattern = "aln/{sample}.bam", variables = { sample = "config.samples_list" } }
]
output = ["merged.bam"]
shell = "samtools merge {output[0]} {input}"   # aln/S1.bam aln/S2.bam ...

Shell loops over the same list work identically (for f in {input}; do ...; done — the pattern used by the multiqc aggregation rules across the community catalog). Note the deliberate split: config.samples_list is comma-joined for expansion-time fan-out, while {input} is space-joined for shell consumption. Do not shell-loop over {config.samples_list} directly — a comma-joined value iterates as a single word; use tr ',' ' ' only if the expansion form is genuinely unavailable. Chunk-level aggregation is the transform operator's combine stage instead (see Transform Operator).

Namespace collision with [[values]] tables — a bare {name} in an expand_inputs pattern that matches a [[values]] table name is treated as that table: the rule fans out once per table entry, and the pattern is expanded entry-by-entry ({chr} with values = ["chr1", …, "chrM"] means 24 separate instances, not "all chromosomes at once"). Writing an aggregation rule that reuses a table name for its literal meaning ("collect every chr's calls") therefore silently produces a per-entry scatter whose instances usually share one output path and race — the repair is a distinct variable name for the aggregation reading, or keying the outputs by the table's wildcard for a genuine per-entry rule:

[[values]]
name = "chr"
values = ["chr1", "chr2", "chr3"]

[config]
chrom_list = ["chr1", "chr2", "chr3"]   # the aggregation binding's source

# WRONG: 'chr' is the [[values]] table — the rule fans into 3 instances
# per pair, every one writing the SAME vcf.raw/{pair_id}.vcf.gz and racing
[[rules]]
name = "merge_chr_vcfs_bad"
expand_inputs = [{ pattern = "vcf.call/{chr}/{pair_id}.vcf.gz" }]
output = ["vcf.raw/{pair_id}.vcf.gz"]

# RIGHT (aggregation): bind the dimension through `variables` instead —
# the pattern expands to ALL entries inside ONE instance ({input} is the
# space-joined list), so the merge happens once
[[rules]]
name = "merge_chr_vcfs"
expand_inputs = [
  { pattern = "vcf.call/{chrom}/{pair_id}.vcf.gz",
    variables = { chrom = "config.chrom_list" } }
]
output = ["vcf.raw/{pair_id}.vcf.gz"]

# RIGHT (scatter): keep the bare table wildcard, but key the outputs by it
[[rules]]
name = "call_chr"
output = ["vcf.call/{chr}/{pair_id}.vcf.gz"]

The two "RIGHT" forms differ in cardinality: the variables-bound aggregation yields one instance whose {input} lists every chromosome; the bare-wildcard scatter yields one instance per chromosome (and typically a downstream aggregation rule over the keyed outputs). The preflight reports the racing shape as SCI-AGG-RACE, naming the unkeyed fan-out dimension (see oxo-flow dry-run → Scientific Preflight).


Wildcards#

Wildcards enable dynamic, pattern-based pipeline definitions. For a detailed guide on how they are discovered, expanded, and constrained, see the Wildcards Reference.

Basic Syntax#

Use {name} in file paths for dynamic expansion:

input = ["raw/{sample}.fastq.gz"]
output = ["aligned/{sample}.bam"]

Built-in Placeholders#

Built-in placeholders use the same syntax but have reserved meanings:

Placeholder Expands to
{input} Space-separated list of all input files
{input[N]} The Nth input file (0-indexed)
{input.name} The value of the name entry in the rule's input map (input = { name = "..." } → {input.name}); any key works: input = { reads = "..." } → {input.reads}
{output} Space-separated list of all output files
{output[N]} The Nth output file (0-indexed)
{output.name} The value of the name entry in the rule's output map (output = { name = "..." } → {output.name}); any key works: output = { bam = "..." } → {output.bam}
{threads} Thread count assigned to this rule
{memory} Memory allocation assigned to this rule
{effective_threads} Declared threads clamped to the machine's CPUs — the tool-facing concurrency. Use it for flags like --threads/-t when the declared value is an HPC-scale label (e.g. a rule declaring threads = 12 renders 4 on a 4-core box). Unset threads render as 1
{effective_memory_mb} Declared memory clamped to the machine's total memory, as an integer number of MB — the tool-facing heap budget. Use it for JVM flags (-Xmx{effective_memory_mb}m, -Xms…) so tools size themselves to the box instead of OOM-ing on hardcoded HPC values (-Xmx29491M class). Unset memory renders as the machine's full total
{log} Path of the rule's log field (every instance wildcard — {sample}, {pair_id}, {assembler} … — and {config.x} are expanded per instance; the parent directory is created automatically)
{config.*} Value from the [config] section (plain value, declared default, or CLI override)
{wildcard.*} Per-instance binding under the same grammar as when conditions — sample-group metadata keys (e.g. {wildcard.umi1_pattern}) and expansion keys. Unbound keys render as their literal brace token

Why the effective pair exists — {threads}/{memory} render the declared values, which keep the pool semantics (an over-capacity rule reserves the whole pool and runs alone). The {effective_*} pair renders what the tool can actually get on this machine. Rules that embed resource-sized flags should use the effective pair; rules that only pass the pool declaration through can keep the plain pair.

Container memory follows the same policy — docker/singularity rules get the declared memory as their --memory cgroup limit, clamped to the machine's total the same way (a rule declaring memory = "72G" on a 4G box runs with --memory 4G, not a limit the kernel can never back). cgroup-aware tools (picard, STAR, Cell Ranger) therefore see an honest limit on every machine size.

The ceiling counts swap — the machine total above is RAM + swap (swap is backable memory the kernel uses under pressure; ignoring it left real capacity unused on small boxes). Pass --max-memory to pin the ceiling to RAM only when latency matters more than headroom.

Existing environments are verified before setup — when the environment cache is cold (e.g. after --rerun wipes the checkpoint) but the conda environment already exists, oxo-flow verifies it in place and marks it ready instead of re-running the setup, whose conda env update fallback re-resolves every dependency over the network.

Environment creation is serialized per environment, not per machine — concurrent processes creating the same env would corrupt each other's transaction, so setup holds an exclusive lock keyed by the env cache key (~/.oxo-flow/locks/env-<sha256>.lock); different envs are independent by conda/pixi semantics and never contend. The wait is bounded (2h by default, OXO_ENV_LOCK_TIMEOUT_SECS to tune): the waiter logs every 60s and a timeout fails the rule with the lock path in the diagnostic instead of hanging the run. The holder's pid is written into the lock file for diagnosis. On shared accounts, one user's slow solve only blocks creators of that same environment.

Spawn watchdog — a rule can in principle sit in the Running state without ever spawning a child process (a lost scheduler wakeup, an environment still being set up, or waiting on the resource pool). Every 30s the engine compares rules it considers running against the processes it can actually see; a rule that has been Running for over OXO_FLOW_SPAWN_WATCHDOG_SECS (default 600) with no child process at all is reported by name, and the report distinguishes two cases (issue #847). If the rule is queued in the resource pool — it has not acquired its threads/memory yet — the report is a one-shot WARN pointing at the scheduler's own 60-second blockage lines, because a queued wait is benign and self-resolving. A rule that is not queued (has acquired resources, or is parked in environment setup) and still has no child process keeps the tracing::error! report — that is the genuine lost-wakeup signature (issue #685). Both reports land in the run log as well as stderr. The watchdog only logs — it never kills; if a named rule never starts, capture the run log and file an issue. Relatedly, the in-process environment-setup mutex now waits on the same bounded schedule as the cross-process lock (OXO_ENV_LOCK_TIMEOUT_SECS, logging every 60s) so a task stuck behind another task's env creation becomes visible instead of hanging silently.

Array-valued {config.*} — a [config] key holding an array renders as a space-joined list in the shell (["a", "b"] → a b), matching the {input} convention; iterate with for x in {config.tools}.

{input} vs {input[N]} — both forms are equivalent when the array has a single entry ({input} joins all inputs with spaces; {input[0]} takes the first). Practical guidance:

  • Single input/output: either form works — {input} and {output} are the simplest
  • Multiple inputs/outputs: use {input[0]}, {input[1]} … to select specific files, or {input} to pass all of them at once
  • The indexed form communicates intent ("this rule expects exactly this file"), so it remains useful even for single-element arrays in multi-step pipelines

An index at or beyond the declared count is never substituted: rendering leaves the literal {input[N]} text in the command and the tool fails with an unrelated "No such file" error. oxo-flow lint catches this statically as W034 (which also flags map-form rules whose sorted-key indexing has shifted under a declaration edit — the index positions are the sorted keys, not the text order).

Named Input & Output#

For complex rules with many files, give the input/output fields an inline map (FilePatterns::Map) and reference the entries with {input.<name>} / {output.<name>}:

[[rules]]
name = "align"
input = { reads1 = "raw/{sample}_R1.fastq.gz", reads2 = "raw/{sample}_R2.fastq.gz" }
output = { bam = "aligned/{sample}.bam" }
shell = "bwa mem {input.reads1} {input.reads2} > {output.bam}"

There is no [rules.named_input] / [rules.named_output] TOML table — the map is the value of input/output itself (a sub-table named named_input fails validation with E017).

A named reference to a key the rule does not declare is likewise never substituted (literal brace text at render time) and is flagged by W034 along with the rule's actual keys, so a typo like {input.tumer} is reported at lint time instead of surfacing as a tool failure.

Custom Wildcards#

Any {name} pattern not matching a built-in placeholder is treated as a wildcard. oxo-flow expands these based on:

  1. File discovery: Scanning for matching files in the input path.
  2. Explicit lists: Defined in [[pairs]] or [[sample_groups]].
  3. Parameter lists: Defined in [[values]].

[[values]] — Parameter Fan-out (WC-03)#

[[values]] declares parameter lists that fan rules out the same way sample groups do — the oxo-flow equivalent of an upstream "for each assembler × binner" construct, without hardcoding one rule per combination:

[[values]]
name = "assembler"
values = ["spades", "megahit"]

[[rules]]
name = "assemble"
input = ["reads/{sample}.fq"]
output = ["assemblies/{assembler}/{sample}/contigs.fa"]
shell = "{assembler} -o {output} {input}"
  • Every rule referencing {assembler} (bare form) or {values.assembler} (namespaced form) expands into one instance per value. The reference may live in any trigger field — inputs, outputs, shell, when, expand_inputs patterns, or the rule's output_pattern (issue #296): a pattern-only reference fans the producer out per value with the pattern baked per instance, exactly as it already does for {sample}/pair wildcards.

[[rules]]
name = "assemble"
output_pattern = "results/{assembler}/chunks.txt"   # fans out per value
shell = "scripts/build_all.sh"
- values_from = "config.x" resolves the value list from a config key at expansion time — a comma string ("u1,u2") or array, the same semantics as expand_inputs config references. CLI --arg x=u1,u2 or a profile then drives the fan-out without editing the workflow; takes precedence over a static values when both are set. - Multiple [[values]] tables form a cartesian product with each other and with sample groups / pairs (deterministic order: the first table varies slowest). - Instance names follow the existing convention: assemble_assembler_spades_batch_S1. - Value groups must have unique names and must not collide with the reserved wildcards (sample, pair_id, experiment, control, group, tumor, normal, experiment_type, tumor_type) or the executor placeholders (input, output, log, threads, memory) — a table named input would replace the placeholder in every rule's shell.


[[pairs]] — Experiment-Control Pairing (WC-01)#

[[pairs]] defines experiment-control sample pairs for somatic variant calling and other comparative analyses.

[[pairs]]
pair_id = "CASE_001"
experiment = "EXP_01"
control    = "CTRL_01"

[[pairs]]
pair_id = "CASE_002"
experiment = "EXP_02"
control    = "CTRL_02"
Field Type Required Description
pair_id String Yes Unique identifier for this pair
experiment String Yes Experiment sample name (alias: tumor)
control String No Matched control sample name (alias: normal)
experiment_type String No Optional cohort label (alias: tumor_type)
metadata Table No Arbitrary key-value pairs (each key becomes a wildcard)
when String No Pair-level gate evaluated at plan time: while it evaluates false, the pair declares no rule instances. Uses the same condition vocabulary as rule when (config.<key> truthiness/comparisons, &&/\|\|/!, parentheses, file_exists(...)). A config.<key> reference with no matching [config] key evaluates false (the unbound→false stance) — a plan-time warning flags the referenced key so a typo does not silently drop the pair's whole rule set

Any rule that references {experiment}, {control}, or {pair_id} in its input, output, or shell fields is automatically expanded into one concrete rule instance per pair. Rules that do not reference any pair wildcard are kept as-is.

Expanded rule naming: {rule_name}_{pair_id} (e.g., mutect2_CASE_001).

Pair-level gating with when#

One static [[pairs]] table can carry profile-switched sample sets — a toggle in [config] decides which pairs fan out:

[config]
extended = false   # flip via --profile or CLI override

[[pairs]]
pair_id = "EXTRA_001"
experiment = "EXP_02"
control = "CTRL_02"
when = "config.extended"

While extended is false, EXTRA_001 fans out no rules; flipping it and re-running restores the pair's instances (re-expansion from the preserved templates, the checkpoint re-entry pattern). Unknown keys evaluate false, so a misspelled key (when = "config.extened") gates the pair out with a plan-time warning naming the offending key.

Loading pairs from external file#

For large cohort studies with hundreds or thousands of pairs, use pairs_file in [workflow]:

[workflow]
name = "somatic-calling"
pairs_file = "metadata/pairs.tsv"  # or .csv, .json

TSV format (tab-separated, header required):

pair_id    experiment    control    experiment_type
CASE_001   EXP_01        CTRL_01    lung_adenocarcinoma
CASE_002   EXP_02        CTRL_02    colorectal
CASE_003   EXP_03        CTRL_03    breast_cancer

CSV format (comma-separated):

pair_id,experiment,control,experiment_type
CASE_001,EXP_01,CTRL_01,lung_adenocarcinoma
CASE_002,EXP_02,CTRL_02,colorectal

JSON format:

[
  {"pair_id": "CASE_001", "experiment": "EXP_01", "control": "CTRL_01"},
  {"pair_id": "CASE_002", "experiment": "EXP_02", "control": "CTRL_02"}
]

Inline [[pairs]] and pairs_file can be used together; entries from both sources are merged. Each pair_id must still come from exactly one source: a pair_id declared in both with identical content is deduplicated with a warning, while the same pair_id with different experiment/control content is a config error naming both declarations (both copies would fan out rules under the same {rule}_{pair_id} name). The same rule applies to duplicate sample-group names across [[sample_groups]] and sample_groups_file.

Auto-discovery from file pattern#

For workflows with existing paired files, use pairs_pattern in [workflow] to auto-discover pairs by scanning the filesystem:

[workflow]
name = "somatic-calling"
pairs_pattern = "aligned/{pair_id}/{experiment}_vs_{control}.bam"

oxo-flow scans files matching this pattern and extracts wildcards from paths. For a file:

aligned/CASE_001/EXP_01_vs_CTRL_01.bam

Creates pair:

  • pair_id = CASE_001
  • experiment = EXP_01
  • control = CTRL_01

Pattern requirements:

  • Must contain {pair_id}, {experiment}, and {control} wildcards
  • Optional {experiment_type} wildcard also extracted
  • Pattern is converted to glob (*) for filesystem scan

This eliminates the need for manual pair lists or external files when working with pre-organized directory structures.

Example#

[[pairs]]
pair_id = "CASE_001"
experiment = "EXP_01"
control    = "CTRL_01"

[[rules]]
name   = "mutect2"
input  = ["aligned/{experiment}.bam", "aligned/{control}.bam"]
output = ["variants/{pair_id}.vcf.gz"]
shell  = "gatk Mutect2 -I {input[0]} -I {input[1]} -normal {control} -O {output[0]}"

Produces rule mutect2_CASE_001 with concrete file paths.

See examples/gallery/15_paired_experiment_control_pairs.oxoflow for a full clinical somatic calling pipeline.


[[sample_groups]] — Multi-Sample Cohorts (WC-02)#

[[sample_groups]] organises samples into named groups (e.g., case vs. control) for cohort studies.

[[sample_groups]]
name    = "control"
samples = ["CTRL_001", "CTRL_002", "CTRL_003"]

[[sample_groups]]
name    = "case"
samples = ["CASE_001", "CASE_002"]
Field Type Required Description
name String Yes Group name
samples Array of strings Yes Sample identifiers in this group
metadata Table No Arbitrary group-level metadata; each key is available as a wildcard in paths, shells, and when conditions (wildcard.<key>)

Any rule that references {sample} or {group} is expanded once per (group, sample) pair across all groups.

Expanded rule naming: {rule_name}_{group}_{sample} (e.g., align_control_CTRL_001).

When groups are loaded from a TSV/CSV sheet, extra columns become metadata keys automatically — see Sheet columns as group metadata.

Loading groups from external file#

For large cohorts, use sample_groups_file in [workflow]:

[workflow]
name = "cohort-analysis"
sample_groups_file = "metadata/groups.tsv"  # or .csv, .json

TSV format (samples can be comma-separated within the field):

name    samples reads_layout    long_reads
short   CTRL_001,CTRL_002   pe  
long    LR_001  ont reads/flowcell1.fq.gz;reads/flowcell2.fq.gz

Every row must supply a value for every column (tab-separated, one row per line) — a row missing a trailing cell fails the sheet parse with an unequal-lengths error. Leave a cell's value empty only if no when gate or placeholder references that metadata key for the group's samples (an empty cell behaves as an unbound key, which closes gates in both directions — including !=).

JSON format:

[
  {"name": "control", "samples": ["CTRL_001", "CTRL_002"]},
  {"name": "case", "samples": ["CASE_001", "CASE_002"]}
]

Like pairs, a group name declared in both [[sample_groups]] and sample_groups_file is deduplicated when identical (with a warning) and rejected as a config error when the samples differ — see the merge rule in Loading pairs from external file.

Sheet columns as group metadata#

In TSV/CSV sheets, every column other than name and samples becomes group metadata: it fans out per (group, sample) pair as a bare {key} wildcard in paths/shell/log and as wildcard.<key> in when conditions — the same vocabulary as the inline metadata table. This is how a cohort declares a second read dimension (issue #283):

[[rules]]
name   = "assemble"
input  = ["{long_reads}"]
output = ["asm/{group}_{sample}/assembly.fa"]
when   = "wildcard.reads_layout == 'ont' || wildcard.reads_layout == 'pacbio'"
shell  = "flye --nano-raw {input} --out-dir {output[0]}"

[[rules]]
name   = "qc"
input  = ["raw/{sample}_R1.fastq.gz"]
output = ["qc/{sample}.txt"]
when   = "wildcard.reads_layout != 'ont' && wildcard.reads_layout != 'pacbio'"
shell  = "fastqc {input} -o {output[0]}"

Short-read rows gate assemble closed and qc open; long-read rows the reverse — one rule template per branch, expanded per sample.

reads_layout vocabulary — the column is validated at plan time; accepted values are pe, single, interleaved (short reads) and ont, pacbio (long reads). Any other value fails the workflow with the accepted set named, because a typo would silently close every when gate built on it. Other metadata columns stay free-form — use the metadata_file table for per-sample free-form values.

Multi-file cells — a metadata cell may carry several paths (e.g. one FASTQ per flowcell) separated by ; (or tab-separated list cells in JSON sheets). They are normalized to a single space-joined value that splices verbatim into the input list as one element; shell rendering joins input-list elements, so flye --nano-raw {input} receives all paths as separate shell tokens (the oxo-flow-varlo target_regions multi-file pattern). Spaces alone do not split cells — separate multiple paths with ;.

Reserved column names — group and sample cannot be used as sheet columns: the fan-out binds {group}/{sample} first and metadata keys last, so such a column would shadow the engine's binding for every row in the group. A sheet carrying them is rejected at parse time — rename the column (e.g. sample_id).

Example#

[[sample_groups]]
name    = "treatment"
samples = ["S001", "S002"]

[[rules]]
name   = "align"
input  = ["raw/{sample}_R1.fq.gz"]
output = ["aligned/{sample}.bam"]
shell  = "bwa mem ref.fa {input[0]} > {output[0]}"

Produces align_treatment_S001 and align_treatment_S002.

See examples/gallery/12_cohort_analysis.oxoflow for a complete cohort study pipeline.

Combining Groups and Pairs — duplicate sample ownership#

A sample may appear in both a user-declared [[sample_groups]] entry and an [[pairs]] entry (e.g. every group sample is also paired against a control). This is legal — group-wildcard rules and pair-wildcard rules expand independently — but the engine warns at parse time, because unless every rule path distinguishes the instances (e.g. a {group} component), the instances share an output path and one is skipped as up-to-date while consuming the other's data. The warning prints once per parse naming all duplicates, not once per duplicate sample:

WARN oxo_flow_core::config::parse: 3 samples are declared by more than one owner: S1 (group 'case' / pair 'P1'), S2 (group 'case' / pair 'P2'), S3 (group 'case' / pair 'P3') — unless every rule path distinguishes them (e.g. a {group} component), the instances share an output path and one is skipped as up-to-date while consuming the other's data

A single duplicate renders a per-sample message instead: sample 'S1' is declared by both group 'case' and pair 'P1' — {same note}. Auto-discovered samples (sample_pattern) feeding [[pairs]] are the documented pairing workflow, not a competing owner — only user-declared groups are compared against pair membership. Fix by adding a {group}/{pair_id} component to the affected rules' output paths, or by keeping the domains disjoint.

Auto-Discovery with sample_pattern#

The [workflow] sample_pattern auto-discovers samples from the filesystem:

[workflow]
# Exactly one pattern — uncomment the one that matches your data:
# Paired-end reads (Illumina standard)
sample_pattern = "raw/{sample}_R1.fastq.gz"

# # Paired-end reads (common variant)
# sample_pattern = "raw/{sample}_1.fq.gz"

# # Single-end reads
# sample_pattern = "raw/{sample}.fastq.gz"

sample_pattern binds {sample} only — {config.*} placeholders are expanded, but any other wildcard name ({read}, {replicate}, {lane}, …) is captured by the glob and never bound: every matching file collapses to the same {sample} value and expansion fails with duplicate rule name. Point the pattern at one file per sample (R1 for paired-end data) and list the mate in each rule's input:

sample_pattern = "raw/{sample}_R1.fastq.gz"

[[rules]]
name  = "align"
input = ["raw/{sample}_R1.fastq.gz", "raw/{sample}_R2.fastq.gz"]
output = ["aligned/{sample}.bam"]
shell = "bwa mem ref.fa {input[0]} {input[1]} > {output[0]}"

Supported wildcards in sample_pattern:

Wildcard Description Example match
{sample} Sample identifier SAMPLE_01 from SAMPLE_01_R1.fastq.gz

For genuine per-read/per-replicate fan-out, declare the domain explicitly with [[sample_groups]], [[pairs]], or [[values]], and reference the extra paths in each rule; input_groups binds additional wildcards discovered from files.

Merging Multiple Sample Sources#

All sample sources are merged into a single {config.samples_list}:

# Filesystem auto-discovery
sample_pattern = "raw/{sample}_R1.fastq.gz"

# CSV/TSV file
sample_groups_file = "metadata/samples.csv"

# Ad-hoc via CLI — filters the declared set (unknown names fail);
# declares samples only when the workflow ships none
oxo-flow run pipeline.oxoflow --samples EXTRA_01,EXTRA_02

All sources deduplicate — the same sample from multiple sources appears once. Per-group sample lists are available as {config.samples_<group_name>}.

These injected values are comma-joined strings (e.g. "S001,S002"):

  • In expand_inputs, scatter.values_from, and split.values_from they resolve per value (comma-split) — see Expand Inputs.
  • In shell templates, {config.samples_list} renders as the comma-joined text, e.g. for s in $(echo {config.samples_list} | tr ',' ' ').

Sample Metadata (metadata_file)#

Per-sample metadata is loaded at plan time from an external TSV/CSV/JSON table and addressed as {meta.<column>} in any rule's input, output, shell, log, script, hooks, and when. This is the metadata group/pair-key mechanism generalized to an arbitrary per-sample column vocabulary — the methylseq when = "config.single_end_mode || {meta.endedness} == 'SE'" pattern, chipseq antibody fan-out (see Input Groups), and rnaseq-style per-sample parameter lookups:

[workflow]
name = "methylseq"
metadata_file = "metadata/samples.tsv"  # or .csv, .json

TSV/CSV format — the first column is the sample id (must match the values your {sample} wildcards resolve to); every other header becomes a column:

sample    endedness    adapters
SE1       SE           AGATCGGAAGAGCACACGTCTGAACTCCAGTCA
PE1       PE           AGATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT

JSON format — a list of objects; the sample key identifies the row (all other keys become columns):

[
  {"sample": "SE1", "endedness": "SE", "adapters": "AGATCGGAAGAGCACACGTCTGAACTCCAGTCA"},
  {"sample": "PE1", "endedness": "PE", "adapters": "AGATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT"}
]

Semantics:

  • The table is a lookup only — it does not define the sample set. Use sample_groups / sample_groups_file / sample_pattern for that.
  • At expansion, each rule instance resolves its metadata row from its {sample} binding, falling back to {pair_id}, then {experiment}, then {control} (first binding that has a row wins).
  • A missing row or column renders empty in shells/inputs/outputs, with a plan-time warning for a column that no row defines (the same stance as {values.name}).
  • In when, metadata values are baked as quoted literals in comparisons ({meta.endedness} == 'SE' → 'SE' == 'SE') and true/false in truthiness positions; a missing row or column renders '' / false — the gate is closed, so samples without the data evaluate false instead of running.
  • A {meta.<column>} reference in a rule that never fans out (no sample-like binding) is left as a residual placeholder — the engine logs a warning at execution time and the literal text remains in the rendered command.
[[rules]]
name   = "trim"
input  = ["raw/{sample}_R1.fastq.gz"]
output = ["trimmed/{sample}.fq"]
when   = "config.single_end_mode || {meta.endedness} == 'SE'"
shell  = "cutadapt -a {meta.adapters} {input[0]} -o {output[0]}"

Metadata paths can also feed inputs directly (DAG edges still infer by exact path match, since the paths are plan-time literals):

[[rules]]
name  = "mapping_qc"
input = ["{meta.bam_path}"]
shell = "samtools stats {input[0]} > {output[0]}"

Partial Pair Tolerance & Tumor-Only Mode#

Pairs with missing controls are supported — control is now optional:

# Tumor-only CNV calling (no matched normals)
[[pairs]]
pair_id = "T1"
experiment = "T1"
# control omitted → {control} = ""

# Pooled normal via config values
[[pairs]]
pair_id = "T2"
experiment = "T2"
# control = ""  # empty string also works

Rules using {control} receive an empty string when no control is specified. Use shell conditionals or declarative [config] entries to handle this:

[[rules]]
name = "cnv_detect"
input = ["aligned/{experiment}.bam"]
shell = """
if [ -n "{control}" ]; then
    cnvkit.py batch {input} --normal aligned/{control}.bam
else
    cnvkit.py batch {input} --method cbs  # tumor-only mode
fi
"""

For pooled-normal scenarios, use declarative [config] entries:

[config]
normal_mode = { default = "pooled", help = "matched, pooled, or none" }
pooled_normal = { default = "results/pooled_normal.bam" }

Supported multi-omics pair patterns:

Scenario Pair Configuration Control
Matched tumor-normal experiment = "T1", control = "N1" Required
Unmatched tumor vs pooled experiment = "T1" None
Tumor-only (CNV, somatic) experiment = "T1" None
Paired-end case-control experiment = "CASE", control = "CTRL" Required
Time-series (no control) two entries: experiment = "T0" and experiment = "T6" None

when — Conditional Rule Execution (WF-01)#

The optional when field on a rule contains an expression evaluated against [config] values. When the expression evaluates to false the rule is skipped at execution time (JobStatus::Skipped, "condition evaluated to false") — it remains in the DAG but does not run. Predicates over per-instance bindings (wildcard.<key>, {meta.<column>}) are decided earlier, at plan time: non-matching instances never enter the DAG, so dry-run and validate counts reflect exactly what a real run would execute.

[[rules]]
name  = "fastqc"
when  = "config.run_qc"
input = ["raw/sample_R1.fq.gz"]
output = ["qc/sample_fastqc.html"]
shell = "fastqc {input[0]} -o qc/"

Expression syntax#

Form Example Description
config.<key> config.run_qc Truthy check (true, non-zero, non-empty string)
config.<key> == "value" config.mode == "WGS" String equality
config.<key> != "value" config.mode != "WES" String inequality
config.<key> == true\|false config.skip == false Boolean equality
config.<key> > N config.min_cov >= 20 Numeric comparison (>, >=, <, <=)
len(config.<key>) len(config.gene_sets) > 0 Element count comparison: array items, string characters, or table keys, compared against a whole number (==, !=, >=, <=, >, <). A missing key counts as 0. Bare len(config.<key>) is truthy when non-empty. Note a bare config.<key> array is truthy even when empty — len(...) > 0 is the "list is non-empty" gate. Comparing against a non-numeric literal is a lint warning (W029) and always evaluates false
file_exists("path") file_exists("panel.bed") File existence test. Relative paths resolve against the workflow root (the workflow file's directory), not the process working directory, and every evaluation site (plan time, execution re-check, config-change analysis) resolves the same way. In per-instance predicates (wildcard.<key> / {meta.<column>}) it is evaluated at plan time — the instance set is fixed when the DAG is built, so files produced mid-run cannot gate instances; express those with {meta.*} (plan-time bakeable) or an execution-time config gate instead
reads_count("path") / wc_lines("path") / file_size("path") reads_count('trimmed/{sample}.fq.gz') > config.min_reads Data-dependent gates — count FASTQ records (4-line records, plain or .gz), text lines of the decompressed content (plain or .gz), or file bytes in the run workdir at execution time. See Data-Dependent when Gates
regex_extract("path", "pattern", group?) regex_extract('alignment/{sample}.flagstat', '(\\d+) \\+ \\d+ mapped', 1) > config.min_mapped_reads Data-dependent gate — parse an integer out of a text report's first regex match (issue #312). group defaults to 0 (the whole match — the pattern must then match a bare number); use group 1+ to extract a number embedded in surrounding text (samtools flagstat's mapped count, bcftools-stats record counts). The captured text must parse as an integer; missing file / no match / non-numeric capture fail closed at execution time. Malformed calls (wrong arity, non-numeric group, uncompilable pattern) are flagged as W030
wildcard.<key> == "value" wildcard.control != "" Per-instance pair/group wildcard comparison (pair_id, experiment, control, tumor, normal, experiment_type, tumor_type, group, sample, plus any metadata keys declared on [[pairs]] / [[sample_groups]] and [[values]] table names) — evaluated per expansion combo at DAG build time; non-matching instances never enter the DAG (snakemake-style morphing)
{meta.<column>} == "value" {meta.endedness} == 'SE' Per-instance sample-metadata comparison (see Sample Metadata) — baked and evaluated per instance at plan time; instances whose gate evaluates false never enter the DAG. A missing row or column renders '' in a comparison (a closed gate)
!<expr> !config.skip Logical NOT
<expr> && <expr> config.run_qc && config.min_cov >= 20 Logical AND
<expr> \|\| <expr> config.wgs \|\| config.wes Logical OR
(<expr>) (config.a && config.b) \|\| config.c Grouping

Any other token — a bare identifier like params_mode, or an unknown function call — is a validation error (E018): the evaluator treats unrecognizable atoms as true, so a typo'd condition would silently let the rule run. Reference config keys as config.<key>, or use a quoted literal, a {...} placeholder, or one of the built-in functions.

Data-Dependent when Gates#

The runtime functions reads_count(path), wc_lines(path), file_size(path), and regex_extract(path, pattern, group?) let a when gate inspect files produced earlier in the same run (issues #282/#312) — e.g. skip a variant caller on samples whose trimmed FASTQ has too few records:

[[rules]]
name  = "call_variants"
when  = "reads_count('trimmed/{sample}.fq.gz') > config.min_reads"
input = ["trimmed/{sample}.fq.gz"]
output = ["variants/{sample}.vcf.gz"]
shell = "caller --input {input[0]} --output {output[0]}"
  • Path resolution — relative paths join against the run workdir; absolute paths pass through. {sample} and other wildcards expand from the instance context before the read. When --workdir differs from the workflow directory, workflow-shipped files are anchored automatically: relative env file specs (conda/mamba/pixi/venv_requirements), rule scripts, input_groups scan patterns, and [[references]].source resolve against the workflow directory. {input} / {log} in rule commands stay workdir-relative — data and logs are addressed from the run cwd (matching the freshness gates and the engine's log-dir creation), so point --workdir at the directory holding your data (issue #427).
  • Phase semantics — at plan time a runtime-function atom whose file is missing — or whose file is an output of this run but still holds a previous run's stale value — defers (true, the rule stays in the plan because its producer may not have run yet); at execution time a missing or unreadable file fails closed (false, the rule is skipped). The defer propagates through !, &&, || and parentheses (three-valued logic), so a negated or chained gate is never pruned at plan time just because its target file is not ready. Planning with oxo-flow dry-run therefore never prunes on a gate whose target file does not exist yet.
  • Formats — reads_count counts 4-line FASTQ records in plain or gzip files (the record count is truncated — a trailing partial record does not round up); wc_lines streams lines of the decompressed content for plain or gzip files (same .gz detection as reads_count); file_size returns the byte length. BAM/CRAM input to these functions is out of scope for now (it needs a BGZF parser, not line arithmetic).
  • Regex extraction — regex_extract(path, pattern, group?) reads the whole file (plain or .gz, up to 16 MiB) and takes the first regex match; the captured text (group 0 = whole match by default, or an explicit capture group) must parse as an integer. Matching is line-oriented ((?m) is applied automatically), so ^ and $ anchor per line. Commas inside the pattern's {} quantifiers, [] character classes, or quoted arguments do not split arguments. The extracted value compares against literals and config.* thresholds exactly like the counting functions.
  • Change tracking — recorded verdicts participate in config-change replay: changing a threshold referenced by such a gate (e.g. config.min_reads) re-evaluates only the instances whose verdict actually flipped, keeping completed unchanged instances exempt.
  • Checkpoint accounting — verdicts are stored in the checkpoint's when_verdicts map. A gate that evaluated to false is a skip, not a failure: the rule never enters failed_rules (a stale entry is pruned), status reports it under Skipped (when condition false) (skipped_by_when in --json), and report --r-data writes it to metrics.tsv as skipped_by_when (issue #690).
  • Scope — only when strings may call these functions; shell, input, and output cannot read the filesystem at plan time.

Unbound wildcard keys evaluate false#

A wildcard.<key> predicate whose key has no binding for an instance — e.g. a group that declares no input_type metadata while another group does — evaluates to false for every operator (==, !=, >, <, …, and bare truthiness alike). The condition is not met, so that instance never runs. This keeps a rule like when = "wildcard.input_type == 'srr'" scoped to the SRA cohort even when other cohorts omit the key entirely.

Note: this is a behavioral contract, not an oversight. An unbound comparison historically evaluated true, which ran SRA-download rules against literal {accession} placeholders for FASTQ cohorts (live incident). oxo-flow lint flags when references to keys no pair/group can ever bind as a W027 warning — such rules never run.

Per-instance wildcard and {meta.<column>} bindings are baked into a kept rule's when at expansion time, so the execution-time re-check re-evaluates the exact same per-instance verdict; config-only predicates continue to gate at execution as before.

wildcard.* is a when-only vocabulary. The dotted form is not a shell placeholder — the engine's placeholder syntax is bare {name}, so {wildcard.input_type} in a shell renders literally. Group/pair metadata keys reach shells as bare placeholders ({input_type}), baked per instance at plan time when the rule fans out on {group}/{sample}; a rule that does NOT fan out per sample has no binding to bake, so such keys stay literal there (use input_groups or {meta.<column>} for those cases).

{meta.<column>} truthiness and boolean columns#

In a bare truthiness position (when = "{meta.do_qc}") a metadata value bakes to true/false with shell-style truthiness: non-empty and not false and not 0 is true — so "0", "", and "false" close the gate and everything else (including "foo") opens it. For a strict boolean column, use the comparison form instead so any unexpected value falls through to no override rather than opening both branches of a markdup/nomarkdup pair:

# strict per-sample boolean (recommended for true/false columns)
when = "config.mark_duplicates || {meta.mark_duplicates} == 'true'"

A missing row or column renders false in a truthiness position and '' in a comparison — always a closed gate, never an unbound token. This makes {meta.x} == '' the official absence guard idiom: it gates the "column present" branch and closes cleanly for rows that lack the column.

Example#

[config]
run_annotation = true
min_coverage   = 30
mode           = "WGS"

[[rules]]
name = "vep_annotate"
when = 'config.run_annotation && config.min_coverage >= 20'
# ...

[[rules]]
name = "wgs_coverage"
when = 'config.mode == "WGS"'
# ...

See examples/gallery/11_conditional_workflow.oxoflow for a full example.


Dependency Resolution#

Dependencies are inferred automatically: if rule B lists a file in its input that appears in rule A's output, then B depends on A.

[[rules]]
name = "step1"
output = ["intermediate.txt"]
# ...

[[rules]]
name = "step2"
input = ["intermediate.txt"]   # depends on step1
# ...

No explicit dependency declaration is needed.


transform — Unified Scatter-Gather Operator#

The transform operator unifies split → map → combine patterns into a single rule declaration, similar to dplyr's group_by() %>% summarize() or pandas' groupby().apply().

Structure#

[[rules]]
name = "variant_calling"
input = ["aligned/sample.bam"]
output = ["variants/sample.vcf.gz"]

[rules.transform.split]
by = "chr"
values_from = "config.chromosomes"

[rules.transform]
map = "gatk HaplotypeCaller -R {config.reference} -I {input} -L {chr} -O {output}"
cleanup = true

[rules.transform.combine]
shell = "gatk GatherVcfs $(for f in {chunks}; do echo \"-I $f \"; done) -O {output}"

Split Configuration#

Field Type Description
by String Required. Variable name for splitting (e.g., "chr", "sample")
values Array Direct list of split values
values_from String Reference to config variable (e.g., "config.chromosomes")
n String Number of map instances — the split values are the labels "0", "1", ..., "n-1". It does not partition the input data
glob String Glob pattern to find split values from files

Priority: values → values_from → n → glob

Split values are labels, not data partitions

The engine never reads, slices, or chunks the declared input. It repeats the rule once per split value and substitutes that value wherever {split_var} appears. With n = "3" and input = ["in.txt"], the map command runs three times with the same input:

process in.txt > .oxo-flow/chunks/chunk/0.txt
process in.txt > .oxo-flow/chunks/chunk/1.txt
process in.txt > .oxo-flow/chunks/chunk/2.txt

Dividing the work is the map command's job: it must use the split value to select its own slice (e.g. -L {chr} for per-chromosome calling), or the rule's input must be the split value (input = ["{chr}"]), so each instance reads its own file. The engine never slices a file, and a split variable embedded in a longer path (shard_{chr}.txt) is left literal. n = "3" alone does neither.

Combine Configuration#

Field Type Description
shell String Shell command to combine chunks
aggregate Boolean Enable automatic aggregation
method String Aggregation method: "concat" or "json_merge"
header String Header line for concat aggregation

Built-in Variables#

Variable Expands to
{split_var} Current split value (e.g., {chr} → "chr1")
{chunks} Space-separated list of all chunk outputs
{input} Chunk outputs — same as {chunks} (in combine)
{output} Original rule output (in combine)

Modes#

Mode A: Split → Map → Combine

Classic scatter-gather with explicit combine command:

[rules.transform.split]
by = "chr"
values_from = "config.chromosomes"

[rules.transform]
map = "gatk HaplotypeCaller -R {config.reference} -I {input} -L {chr} -O {output}"

[rules.transform.combine]
shell = "gatk GatherVcfs $(for f in {chunks}; do echo \"-I $f \"; done) -O {output}"

Mode B: Split → Map (No Combine)

Parallel processing without merging — each split produces independent output:

[rules.transform.split]
by = "chr"
values_from = "config.chromosomes"

[rules.transform]
map = "samtools flagstat {input} > {output}"
# No combine section

Mode C: Split → Map → Aggregate

Automatic aggregation (concat or json_merge):

[rules.transform.split]
by = "chunk"
n = "5"

[rules.transform]
map = "process {input} > {output}"

[rules.transform.combine]
aggregate = true
method = "concat"

Cleanup#

When cleanup = true, chunk files are automatically deleted after the whole run finishes successfully (empty chunk directories are removed too):

[rules.transform]
cleanup = true

Failed runs keep their chunks for debugging. Because chunk deletion makes the map outputs "missing", a re-run always recomputes the map rules — that is the disk-space trade-off of cleanup = true.

Failure and Retry Logic#

In a scatter-gather process, failures are handled at the chunk (map) level:

  • If a single chunk fails, only that specific chunk is retried according to the rule's retries setting.
  • Sibling chunks continue to process in parallel.
  • The combine step will not execute until all chunks succeed. If any chunk fails exhaustively (after all retries), the combine step is cancelled.

Expanded Rule Naming#

Transform rules expand into:

  • Map rules: {rule_name}_{split_value} (e.g., variant_calling_chr1)
  • Combine rule: {rule_name}_combine (e.g., variant_calling_combine)

Multi-line Strings#

Use triple quotes for multi-line shell commands:

shell = """
mkdir -p results
bwa mem -t {threads} ref.fa {input} | \
  samtools sort -@ {threads} -o {output}
"""

Backslashes inside \"\"\" blocks are TOML escape sequences

"""…""" is a TOML basic string: escape sequences like \n, \t, and \\ are decoded before the shell ever sees the text. Two consequences:

  • The \ line continuations above only work because TOML folds backslash-newline into nothing. But an embedded heredoc or regex that needs a literal \n breaks — the shell receives a real newline instead.
  • For scripts with literal backslashes (embedded python3 <<'EOF' heredocs, sed 's/\t/,/g', …), use a multi-line literal string with triple single quotes — TOML passes its contents through byte-for-byte:
shell = '''
python3 - <<'PYEOF'
print("a\nb")   # literal backslash-n reaches python intact
PYEOF
'''

Complete Example#

[workflow]
name = "ngs-pipeline"
version = "2.0.0"
description = "Complete NGS analysis pipeline"
author = "Shixiang Wang <w_shixiang@163.com>"

[config]
reference = "/data/ref/hg38.fa"
known_sites = "/data/ref/known_sites.vcf.gz"
results = "results"

[defaults]
threads = 4
memory = "8G"
environment = { conda = "envs/base.yaml" }

[report]
format = ["html"]

[[rules]]
name = "fastqc"
input = ["raw/{sample}_R1.fastq.gz", "raw/{sample}_R2.fastq.gz"]
output = ["{config.results}/qc/{sample}_R1_fastqc.html"]
shell = "fastqc {input} -o {config.results}/qc/ -t {threads}"

[[rules]]
name = "trim"
input = ["raw/{sample}_R1.fastq.gz", "raw/{sample}_R2.fastq.gz"]
output = [
    "{config.results}/trimmed/{sample}_R1.fastq.gz",
    "{config.results}/trimmed/{sample}_R2.fastq.gz"
]
environment = { docker = "biocontainers/fastp:0.23.4" }
shell = "fastp --in1 {input[0]} --in2 {input[1]} --out1 {output[0]} --out2 {output[1]} --thread {threads}"

[[rules]]
name = "align"
input = [
    "{config.results}/trimmed/{sample}_R1.fastq.gz",
    "{config.results}/trimmed/{sample}_R2.fastq.gz"
]
output = ["{config.results}/aligned/{sample}.bam"]
environment = { conda = "envs/alignment.yaml" }
shell = "bwa mem -t {threads} -R '@RG\\tID:{sample}\\tSM:{sample}' {config.reference} {input[0]} {input[1]} | samtools sort -o {output}"

[rules.resources]
threads = 16
memory = "32G"

JSON Schema#

oxo-flow ships a JSON Schema for the .oxoflow format. Use it for automated validation in your CI/CD pipelines or for real-time autocompletion and error checking in your IDE (like VS Code or IntelliJ).

Getting the Schema#

You can output the schema directly from the CLI:

oxo-flow schema > oxoflow.schema.json

IDE Configuration (VS Code)#

To enable validation in VS Code, add the following to your settings.json:

"yaml.schemas": {
    "https://traitome.github.io/oxo-flow/schema/oxoflow-v1.schema.json": "*.oxoflow"
}

(Note: Although .oxoflow is TOML, many VS Code extensions can apply JSON schemas to multiple formats).


Checkpoint Re-entry#

A checkpoint = true rule discovers new values at runtime (Snakemake-style checkpoint re-entry): after it completes, the engine reads its checkpoint_manifest — a TOML file the rule itself wrote — and re-expands the workflow with the new values. Every round is still a static plan, so previews stay deterministic and resumes reconstruct the same plan.

[[rules]]
name = "discover"
output = ["discover.toml"]
shell = "python discover.py > discover.toml"
checkpoint = true
checkpoint_manifest = "discover.toml"

The manifest declares new wildcard values — new samples and/or new experiment-control pairs in the same round:

[reentry]
group = "batch"            # optional; defaults to the workflow's first sample group
sample = ["S4", "S5"]      # appended (dedup) to that group
pairs = [                  # optional; appended (dedup by pair_id) to [[pairs]]
  { pair_id = "CASE_007", tumor = "T7", normal = "N7", tumor_type = "tumor" },
]

Pair entries mirror [[pairs]] fields: pair_id (required), experiment (required, alias tumor), control (alias normal), experiment_type (alias tumor_type), and an optional metadata table. tumor_type and metadata are optional.

Semantics:

  • On success, the engine merges the samples and pairs, then re-expands the rule templates; newly created instances (e.g. analyze_batch_S4, call_CASE_007) execute in the same run. Existing instances are untouched.
  • The checkpoint records each re-entry (reentries array: round, rule, group, samples, pairs). A resume replays the records whose checkpoint rule is still up-to-date — the plan reconstructs identically. If the checkpoint rule is invalidated (input/config change), its contribution is revoked (samples and pairs) until it re-runs and re-records.
  • Pair identity is the pair_id: an existing id with identical content is a no-op; an existing id with different content is an error (E015) — silently superseding it would corrupt already-run pair outputs. The same sample appearing in several pairs is not an ambiguity: pair instances are keyed by pair_id, so each pair is its own instance.
  • A discovered pair_id whose instance name collides with an existing instance (e.g. a sample-group instance analyze_CASE_007) is an error (E016).
  • An empty manifest (sample = [], no pairs) is a valid no-op. A missing or unparsable manifest fails the checkpoint rule with a clear error, and its dependents do not run.
  • Bounds: checkpoint rules are never re-expanded themselves (no {sample}/{group}/{pair_id}/{experiment} — validation error E014) and re-entry is capped at 32 rounds — a rule that keeps discovering values past that is a workflow bug, not an engine feature. Diagnostic E013 (emitted at warning severity — checkpoint = true predates the manifest field) flags a checkpoint rule that does not declare checkpoint_manifest.
  • dry-run previews replay recorded re-entries (the preview shows the same static plan a run would execute) and mark checkpoint rules as possible re-entry points; --json includes a reentry section.

Runtime-Discovered Outputs (output_pattern)#

An output_pattern rule declares its outputs by filesystem pattern instead of a fixed output list: after the rule completes, the engine scans the working directory for files matching the pattern, and any downstream rule whose input references the pattern's wildcard is instantiated once per discovered value. This is the "checkpoint-lite" variant of re-entry — the discovered values come from the file system, not from a manifest the rule wrote (which is what makes it work for rules whose file counts are unknown ahead of time, e.g. interval scatters or contig-aware variant calling).

[[rules]]
name = "split"
output_pattern = "results/chunks/{i}.txt"
shell = "for i in 1 2 3; do echo $i > results/chunks/$i.txt; done"

[[rules]]
name = "collect"                       # deferred: no {i} instances at plan time
input = ["results/chunks/{i}.txt"]
output = ["results/merged/{i}.txt"]
shell = "cat {input} > {output}"

split produces results/chunks/1.txt, results/chunks/2.txt, results/chunks/3.txt; the engine instantiates collect_1, collect_2, collect_3 at runtime and inserts them into the plan, so the run completes with all three merged files. The fresh wildcard may combine with existing sample wildcards — a producer fanned out over {sample} contributes its domain per sample instance, and consumers instantiate over the union.

Semantics:

  • output_pattern is mutually exclusive with output and with transform and input_groups (validation errors; v1 exclusions). The pattern needs at least one wildcard — a fixed pattern can never enumerate a domain.
  • {config.x} in the pattern resolves from the workflow config before the scan; discovered files are literals, so exact-match edge inference works as usual.
  • A [[values]] wildcard in the pattern is not fresh — it is a plan-time bound dimension (issue #296). The producer fans out per value with the pattern baked per instance, the pattern is registered producer-side in the DAG (so a per-value consumer's materialized input exact-matches its own producer instance), and a pattern whose wildcards are all bound this way skips the runtime scan entirely — the static instances already are the domain. Mixed patterns (results/{assembler}/{chunk}.txt) compose: the values dimension bakes per instance while the fresh wildcard stays runtime-discovered, and the deferred consumers instantiate over the union with both bindings.
  • Plan time: a consumer referencing a fresh wildcard has an empty domain and instantiates nothing; it is deferred whole. Wildcards shared with the producer's pattern ride the discovered bindings; [[values]] tables the consumer references on its own project as an extra Cartesian dimension at instantiation, under the same plan-time semantics as a static fan-out — wildcard_constraints filtered and per-instance when gates evaluated, so gated-off combos never become instances (issue #296 follow-up). A consumer declared before its producer is legal but warned (output_pattern consumer declared BEFORE its producer; the fan-out still resolves, but declare the producer first for stable run attribution). Referencing two producers' fresh wildcards in one rule is a v1 error; a chain — a rule that is both consumer and producer of fresh wildcards — is supported.
  • Runtime: when a producer instance completes, the engine scans its pattern and unions the discovered values into the producer template's domain. When all of the producer's instances have completed, the deferred consumers are instantiated (one instance per domain value × own values binding, deterministically named consumer_<own values>_<values>, with expand_inputs patterns materialized into concrete inputs), and the new instances are inserted forward into the remaining plan — forward-safety comes from completion ordering, not DAG topology, exactly like checkpoint re-entry.
  • Interaction with resume: the discovered domains are persisted (persist-first, before fan-out). A resume replays the persisted domain and instantiates the deferred consumers only when every producer instance is already completed; when a failed instance still has to re-run, its consumers stay pending and the live runtime pass instantiates them from the final domain union — the final consumer set equals the union of all per-instance discoveries, never a partial slice.
  • Zero discoveries after producer success is loud: a warning is emitted and the consumers stay uninstantiated (a later producer instance may still contribute). A producer failure leaves the consumers uninstantiated and the failure summary attributes them — "not instantiated: producer X failed" — never a silent zero-match.
  • Checkpoints persist the discovered domain (output_pattern_domains), so a resume re-instantiates the consumers without re-running the producer; re-instantiation is idempotent (same names, same plan).

[webhook] — Workflow-Level Notifications#

Post workflow_started / workflow_completed / workflow_failed events to an HTTP endpoint when a run starts and finishes (issue #227 item 1) — the transport counterpart of the on_complete / on_error shell hooks.

[webhook]
url = "https://hooks.example.com/oxo"
# method = "POST"                  # POST (default) | PUT | GET
# events = ["workflow_completed", "workflow_failed"]  # default: workflow_completed
# headers = { Authorization = "Bearer …" }
# secret = "…HMAC key…"            # X-OxoFlow-Signature header (HMAC-SHA256)
# signature_scheme = "hmac-sha256"
# timeout_secs = 30
# max_retries = 3
Field Type Default Description
url String — (required) Webhook endpoint URL
method String POST HTTP method (POST/PUT/GET)
events Array of String ["workflow_completed"] Which events fire: workflow_started, workflow_completed, workflow_failed (rule_completed/rule_failed are defined but not yet dispatched)
headers Table {} Custom request headers
secret String — HMAC key for the X-OxoFlow-Signature header
signature_scheme String hmac-sha256 hmac-sha256 (RFC 2104) or the legacy sha256-keyed
timeout_secs Integer 30 Per-request timeout
max_retries Integer 3 Retries with exponential backoff

The payload is JSON: {event, workflow_name, timestamp, data, version} where data carries the run counters (succeeded, failed, skipped) and, on failure, an error summary. Notifications are best-effort — an unreachable or erroring endpoint logs a warning and never changes the run status or its exit code.

[engine] — Runtime Disk-Pressure Guard (issue #843)#

Optional run-time disk monitoring. Without this section the engine never measures free space mid-run and behavior is unchanged.

[engine]
min_free_disk = "50G"        # required to enable: abort/park floor
# reclaim_free_disk = "100G" # optional: reclaim trigger; defaults to min_free_disk
Field Type Default Description
min_free_disk String — Floor for free disk space at the working directory. A bare number means MiB. Sizes accept T/G/M/K suffixes (binary units: 1G = 1024 MiB). Note "50GB" is not valid — the unit is a single letter
reclaim_free_disk String min_free_disk Reclaim trigger level; must be ≥ min_free_disk (validation error otherwise)

Response ladder#

Each measurement (run start, loop head, and every 60 s while the run is parked or between dispatch batches) classifies the working directory's free space and reacts:

Level Condition Response
Normal free ≥ reclaim Run proceeds normally
Reclaim min ≤ free < reclaim Completed rules' temporary outputs are reclaimed (tombstoned in the checkpoint + unlinked) to make room; the run continues
Hold free < min New rule dispatches are withheld (in-flight rules continue); if nothing is in flight and nothing is reclaimable, the run parks and re-measures every 60 s
Abort Hold persists for ~2 consecutive parked rounds (~2 min) The run checkpoints its state and fails with the disk-pressure message

Under Hold with rules in flight, in-flight rules run to completion (their outputs may themselves trigger Reclaim), but no new rule starts until free space recovers above the floor.

The abort message explains recovery: free space manually (delete intermediates or run oxo-flow clean), then oxo-flow resume from the checkpoint — or lower min_free_disk.

What gets reclaimed#

Reclamation targets a completed rule's outputs only when all of the following hold: the rule is marked temporary = true, it is not a leaf (it has dependents), every dependent has completed and each dependent's own declared outputs still exist on disk — a completed rule whose outputs went missing will re-execute and needs its inputs in place, so reclaiming them underneath it would break the re-run (the reclaim-vs-resume race, fixed under issue #843), the paths are not protected_output, and the files exist on disk. Reclaimed rules get a tombstone in checkpoint.json — the same mechanism the success-path temporary cleanup uses — so a later run that needs the outputs regenerates them first (lazy cascade-up). See Cleanup Behavior.

Notes:

  • Measurements are fail-open: if the working directory cannot be stat'ed, the engine logs a warning and continues rather than aborting.
  • On non-Unix platforms the monitor is inactive (a warning is printed at startup when [engine] thresholds are set).
  • A run that starts already below min_free_disk prints a warning before executing (it does not refuse to start; the Hold/Abort ladder governs from there).
  • The park/abort window is tunable via environment: OXO_FLOW_DISK_TICK_SECS (seconds between measurements while parked, default 60) and OXO_FLOW_DISK_HOLD_ROUNDS (grace park rounds before the abort, default 2). Unset, invalid, or 0 values fall back to the defaults, so a typo can never disable the abort.
  • Free space is measured at the working directory's filesystem (statvfs), which is volume-level: thresholds apply to the whole volume, not to a per-directory quota. Tools that spill large scratch outside the workdir (honor TMPDIR deliberately if needed) are invisible to the monitor.
  • Mid-run reclaim of fan-out producers (output_pattern with per-instance values) does not fire: the event loop tracks reclaim candidates by rule name and holds no instance bindings. Fail-safe — no reclaim rather than a wrong one. The post-run sweep covers them (see below).
  • Only paths declared by the rule (its output entries / matching output_pattern instances) are reclaimed; sidecar files a command writes without declaring are never touched. Declare index sidecars explicitly (e.g. a genome.fa rule also declaring genome.fa.fai) so reclaim + cleanup can track them.

See Also#