Skip to content

Introduction#

oxo-flow is a Rust-native bioinformatics pipeline engine built from first principles for performance, reproducibility, and environment isolation. It compiles workflows into a Directed Acyclic Graph, manages software environments automatically, and runs jobs in parallel — all from a single, fast binary with no external runtime required.

# Define your workflow in TOML
cat pipeline.oxoflow

# Execute it
oxo-flow run pipeline.oxoflow -j 8
oxo-flow v0.23.2 — Rust-native bioinformatics pipeline engine
DAG: 5 rules in execution order
  1. fastqc
  2. trim_reads
  3. bwa_align
  4. sort_bam
  5. call_variants
  Running: fastqc
  ✓ fastqc (3.2s)
  ⋮
Done: 5 succeeded, 0 skipped, 0 failed
✓ 9 output files verified (3.4GB total)

Community catalog

Looking for ready-to-run workflows? The community maintains a curated, rated catalog at oxo-flow-community — verified ports of popular nf-core & Snakemake pipelines, original workflows, and community submissions.


What Is oxo-flow?#

oxo-flow is a high-performance workflow engine built from the ground up in Rust for bioinformatics and clinical genomics. You define pipelines in a clean TOML format (.oxoflow files), and oxo-flow handles dependency resolution, environment activation, parallel execution, and report generation — with compile-time safety guarantees and zero interpreter overhead.

Core capabilities#

Capability Description
DAG engine Automatic dependency resolution, topological sorting, cycle detection, and parallel execution groups
Environment management First-class support for 8 backends: conda, mamba, pixi, docker, singularity, venv, system, HPC modules — per rule
Reporting Generate structured HTML, JSON, Markdown, and PDF reports from checkpoint data — execution truth, failure diagnosis, and file checksums
Web API Built-in REST API (axum-based) for building, validating, and monitoring workflows remotely
VS Code extension First-party traitome.oxo-flow IDE for .oxoflow — schema-driven completion, background diagnostics, canonical formatting, one-click run / dry-run / graph / status / clean / resume / AI-generate (published to Open VSX; also attached as a VSIX to every release)
Container packaging Package entire workflows into Docker or Singularity images for portable, reproducible execution
Cluster backends Submit jobs to SLURM, PBS, SGE, and LSF clusters with resource-aware scheduling
Wildcard expansion {sample}, {chr} patterns that expand automatically from inputs or config

Who Is This For?#

Bioinformaticians who build and maintain analysis pipelines — oxo-flow gives you a faster, type-safe workflow engine with clear error messages, reproducibility guarantees, and no external runtime dependency.

Research laboratories and core facilities running genomics workflows — the reporting system produces structured, auditable reports (engine version, workflow checksum, checkpoint provenance), and container packaging ensures reproducibility across environments.

Researchers who need reproducible science — every workflow execution is deterministic, and environments are locked per rule so results are the same on any machine.

Core facility staff managing multi-sample, multi-assay workloads — the DAG engine and cluster backends handle parallelism and resource scheduling automatically.


How to Use This Guide#

This documentation follows the Diátaxis framework and is organized into four sections:

If you are new to oxo-flow#

Start with the Tutorials in order:

  1. Installation — install the binary
  2. Quick Start — run your first workflow in 5 minutes
  3. Your First Workflow — build a pipeline from scratch
  4. Writing Custom Scripts — embed Python, R, and Bash scripts
  5. Variant Calling Pipeline — complete NGS analysis
  6. Environment Management — conda, mamba, pixi, docker, singularity, venv, system, HPC modules

If you want to learn by example#

Explore the Workflow Gallery — 16 curated workflows from hello-world to multi-omics integration and applied study designs, each with validation output, DAG visualizations, and scientific context:

  1. Hello World ⭐ — Minimal rule structure
  2. File Pipeline ⭐⭐ — Multi-rule dependencies
  3. Parallel Samples ⭐⭐ — Wildcard expansion
  4. Scatter-Gather ⭐⭐⭐ — Parallel chunk processing
  5. Environment Management ⭐⭐⭐ — Per-rule isolation
  6. RNA-seq Quantification ⭐⭐⭐⭐ — Transcriptomics pipeline
  7. WGS Germline Calling ⭐⭐⭐⭐⭐ — GATK best practices
  8. Multi-Omics Integration ⭐⭐⭐⭐⭐ — WGS + RNA-seq + Methylation
  9. Single-Cell RNA-seq ⭐⭐⭐⭐ — scRNA-seq analysis
  10. Transform Operator ⭐⭐⭐ — Unified scatter-gather
  11. Conditional Execution ⭐⭐⭐ — when-gated WGS/WES branches
  12. Cohort Analysis ⭐⭐⭐⭐ — Multi-group cohorts and QC aggregation
  13. Germline Variant Calling ⭐⭐⭐⭐ — Per-sample GATK chain
  14. Paired Experiment-Control ⭐⭐⭐⭐⭐ — Single-pair somatic calling
  15. Paired Experiment-Control (Pairs) ⭐⭐⭐⭐⭐ — [[pairs]]-scaled somatic calling
  16. 16S Amplicon (QIIME2) ⭐⭐⭐⭐ — Amplicon denoising and diversity

If you need to accomplish a specific task#

Jump to the How-to Guides:

If you need exact syntax and options#

See the Command Reference for all 30 CLI subcommands with usage, options, and examples, or the VS Code Extension reference for the first-party editor integration.

If you want the full technical details#

See Architecture & Design for in-depth documentation of the DAG engine, environment system, .oxoflow format specification, and web API.


Quick Example#

Here is a complete workflow that aligns paired-end reads and sorts the output:

# align.oxoflow
[workflow]
name = "align-and-sort"
version = "1.0.0"
sample_pattern = "{sample}_R1.fastq.gz"

[config]
reference = "/data/ref/hg38.fa"

[defaults]
threads = 4
memory = "8G"

[[rules]]
name = "bwa_align"
input = ["{sample}_R1.fastq.gz", "{sample}_R2.fastq.gz"]
output = ["aligned/{sample}.bam"]
environment = { docker = "quay.io/biocontainers/bwa:0.7.19--h577a1d6_1" }
shell = "bwa mem -t {threads} {config.reference} {input} | samtools sort -o {output}"

[rules.resources]
threads = 16
memory = "32G"

[[rules]]
name = "index_bam"
input = ["aligned/{sample}.bam"]
output = ["aligned/{sample}.bam.bai"]
environment = { conda = "envs/samtools.yaml" }
shell = "samtools index {input}"

{sample} is a wildcard: sample_pattern tells oxo-flow to discover one sample per matching read file (S1_R1.fastq.gz, S2_R1.fastq.gz, … become one bwa_align task per sample). See Wildcards for the other expansion sources.

Run it:

# Validate the workflow
oxo-flow validate align.oxoflow

# Preview the execution plan
oxo-flow dry-run align.oxoflow

# Execute (dry-run prints a suggested -j; the resource pool also
# schedules by thread capacity, so jobs queue instead of oversubscribing)
oxo-flow run align.oxoflow -j 1

# Visualize the DAG (write the dot description, then render it)
oxo-flow graph -f dot -o dag.dot align.oxoflow
dot -Tpng -o dag.png dag.dot

stdout is for machines, stderr is for humans

Progress, warnings, and log lines go to stderr; stdout is reserved for machine-readable output — --json, graph --format dot, and every --format/-o deliverable. That is why the graph example above writes the DOT to a file with -o (or redirect stdout) instead of mixing it with run logs.


Key Concepts#

If you are new to pipeline engines, here are the three core concepts used in oxo-flow:

Workflow#

A Workflow is the entire pipeline definition (usually a .oxoflow file). It contains a collection of rules, configuration settings, and software requirements.

Rule#

A Rule is a single processing step. It defines:

  • Input: The files needed to run (e.g., raw reads).
  • Output: The files produced by the step (e.g., aligned BAM).
  • Command: The actual shell command to execute (e.g., bwa mem).
  • Environment: The software tools needed (e.g., a specific Conda environment).

DAG (Directed Acyclic Graph)#

A DAG is a mathematical representation of your workflow. It is a "map" that shows how rules are connected by their inputs and outputs. oxo-flow builds this map automatically to determine which rules can run in parallel and which must wait for others to finish.


Project Status#

oxo-flow is under active development. The current release (v0.23.2) includes the complete core engine, CLI, and web API. See the Changelog for release history and the Contributing guide if you want to get involved.


How to Cite#

If you use oxo-flow in academic research, please cite:

Shixiang Wang, oxo-flow: compiled, memory-safe bioinformatics workflow orchestration, bioRxiv, 2026, https://doi.org/10.64898/2026.06.11.731578

Jia Ding, Yun Peng, Ruochen Wei, Boquan Wang, Jian-Guo Zhou, Shixiang Wang, BLIT: an R package for seamless integration of command-line bioinformatics tool universe, Bioinformatics Advances, Volume 6, Issue 1, 2026, vbag088, https://doi.org/10.1093/bioadv/vbag088

Join the Community#

oxo-flow is a community-driven, open-source project licensed under Apache 2.0. Bug reports, feature requests, and contributions are welcome.

How to contribute Link
Report a bug Bug report
Request a feature Feature request
Contribute code Contributing guide

Try it, break it, and tell us what happened. Even a short comment about what worked — or didn't — helps improve oxo-flow for everyone.