Skip to content

Your First Workflow#

This tutorial walks you through building a realistic bioinformatics workflow from scratch. You will create a quality-control pipeline that processes FASTQ files through FastQC and fastp, then generates a summary report.


Prerequisites#

  • oxo-flow installed
  • Paired-end FASTQ files (or the willingness to create test files)
  • conda or mamba available for environment management.
    • If you don't have either, we recommend Miniforge3, which includes both conda and mamba.

1. Set up the project#

oxo-flow init qc-pipeline
cd qc-pipeline

2. Create environment files#

Create a conda environment file for the QC tools:

# envs/qc.yaml
name: qc
channels:
  - bioconda
  - conda-forge
dependencies:
  - fastqc=0.12.1
  - fastp=1.3.6
  - multiqc=1.35

3. Write the workflow#

Replace qc-pipeline.oxoflow with:

Configuration Syntax

{config.samples_dir} refers to the samples_dir variable defined in the [config] section. This allows you to centralize paths and settings.

Wildcard Patterns

The {sample} in the file paths below is a wildcard. To expand it, declare a sample_pattern in the [workflow] section: oxo-flow scans your raw_data directory for files matching the pattern {sample}_R1.fastq.gz, extracts the sample name, and automatically generates a task for every sample it finds.

[workflow]
name = "qc-pipeline"
version = "1.0.0"
description = "Quality control for paired-end sequencing data"
author = "Your Name"
sample_pattern = "raw_data/{sample}_R1.fastq.gz"   # ← {config.*} is expanded here too: "{config.samples_dir}/{sample}_R1.fastq.gz" works

[config]
samples_dir = "raw_data"
results_dir = "results"

[defaults]
threads = 4
memory = "8G"

[[rules]]
name = "fastqc_raw"
input = [
    "{config.samples_dir}/{sample}_R1.fastq.gz",
    "{config.samples_dir}/{sample}_R2.fastq.gz"
]
output = [
    "{config.results_dir}/fastqc/{sample}_R1_fastqc.html",
    "{config.results_dir}/fastqc/{sample}_R1_fastqc.zip",
    "{config.results_dir}/fastqc/{sample}_R2_fastqc.html",
    "{config.results_dir}/fastqc/{sample}_R2_fastqc.zip"
]
environment = { conda = "envs/qc.yaml" }
shell = """
mkdir -p {config.results_dir}/fastqc
fastqc {input} -o {config.results_dir}/fastqc -t {threads}
"""

[[rules]]
name = "fastp_trim"
input = [
    "{config.samples_dir}/{sample}_R1.fastq.gz",
    "{config.samples_dir}/{sample}_R2.fastq.gz"
]
output = [
    "{config.results_dir}/trimmed/{sample}_R1.fastq.gz",
    "{config.results_dir}/trimmed/{sample}_R2.fastq.gz",
    "{config.results_dir}/trimmed/{sample}_fastp.html",
    "{config.results_dir}/trimmed/{sample}_fastp.json"
]
environment = { conda = "envs/qc.yaml" }
shell = """
mkdir -p {config.results_dir}/trimmed
fastp \
  --in1 {config.samples_dir}/{sample}_R1.fastq.gz \
  --in2 {config.samples_dir}/{sample}_R2.fastq.gz \
  --out1 {config.results_dir}/trimmed/{sample}_R1.fastq.gz \
  --out2 {config.results_dir}/trimmed/{sample}_R2.fastq.gz \
  --html {config.results_dir}/trimmed/{sample}_fastp.html \
  --json {config.results_dir}/trimmed/{sample}_fastp.json \
  --thread {threads}
"""

[[rules]]
name = "fastqc_trimmed"
input = [
    "{config.results_dir}/trimmed/{sample}_R1.fastq.gz",
    "{config.results_dir}/trimmed/{sample}_R2.fastq.gz"
]
output = [
    "{config.results_dir}/fastqc_trimmed/{sample}_R1_fastqc.html",
    "{config.results_dir}/fastqc_trimmed/{sample}_R1_fastqc.zip"
]
environment = { conda = "envs/qc.yaml" }
shell = """
mkdir -p {config.results_dir}/fastqc_trimmed
fastqc {input} -o {config.results_dir}/fastqc_trimmed -t {threads}
"""

[[rules]]
name = "multiqc"
input = [
    "{config.results_dir}/fastqc/sample1_R1_fastqc.html",
    "{config.results_dir}/fastqc_trimmed/sample1_R1_fastqc.html"
]
output = [
    "{config.results_dir}/multiqc/multiqc_report.html"
]
threads = 1  # override [defaults] threads = 4 — aggregation is I/O-light
environment = { conda = "envs/qc.yaml" }
shell = """
mkdir -p {config.results_dir}/multiqc
multiqc {config.results_dir} -o {config.results_dir}/multiqc --force
"""

4. Understand the dependency graph#

The workflow forms this DAG:

graph TD
    A[fastqc_raw] --> D[multiqc]
    B[fastp_trim] --> C[fastqc_trimmed]
    C --> D
  • fastqc_raw and fastp_trim can run in parallel (no dependency between them)
  • fastqc_trimmed depends on fastp_trim's output — inferred automatically because its input files match fastp_trim's output files
  • multiqc aggregates the two QC rounds — its two inputs are the sample1 report files produced by fastqc_raw and fastqc_trimmed. Note that fastp_trim has no direct edge to multiqc: the trimmed data itself is not a multiqc input; multiqc implicitly waits for it via fastqc_trimmed's transitive dependency. Because neither the inputs nor the output of multiqc contain {sample}, it stays a single task and waits for both per-sample QC branches.

Parallel scheduling

The engine does not serialize the whole pipeline. fastqc_raw starts immediately, in parallel with fastp_trim — it never waits for trimmed data. Only fastqc_trimmed waits for fastp_trim to finish, and multiqc waits for both QC rounds. On a multi-core machine, raw QC and trimming run simultaneously.

Two dependency mechanisms

oxo-flow supports two ways to declare dependencies:

  1. File-based (automatic) — if rule B's input matches rule A's output, the edge is inferred. This tutorial uses only this mechanism: fastp_trim → fastqc_trimmed (trimmed reads) and fastqc_raw / fastqc_trimmed → multiqc (QC reports).
  2. depends_on (explicit) — list rule names that must finish first, even when no direct file match exists. Use this for setup rules with no outputs:
[[rules]]
name = "setup_dirs"
output = []                  # no files to match!
shell = "mkdir -p results"

[[rules]]
name = "align"
depends_on = ["setup_dirs"]  # ← explicit ordering, no file to match
shell = "bwa mem ..."

Prefer file-based inference when possible — it makes the data flow self-documenting. Use depends_on only when the ordering can't be expressed through file matching alone.


5. Prepare Test Data#

For this tutorial, create minimal test files so oxo-flow has something to process. They must be parseable FASTQ — 4 lines per record with the sequence and quality lines the same length — or FastQC/fastp abort with a SequenceFormatException instead of producing their reports:

mkdir -p raw_data
# 10 synthetic 60 bp reads per file (varied sequence, Q40 qualities).
for sample in sample1 sample2; do
  for read in 1 2; do
    awk -v sample="$sample" 'BEGIN {
      bases = "ACGT"
      for (i = 1; i <= 10; i++) {
        seq = ""
        for (j = 0; j < 60; j++) seq = seq substr(bases, (i + j) % 4 + 1, 1)
        qual = ""
        for (j = 0; j < 60; j++) qual = qual "I"
        print "@" sample ":" i
        print seq
        print "+"
        print qual
      }
    }' | gzip > "raw_data/${sample}_R${read}.fastq.gz"
  done
done

Each file ends up as 10 valid records, e.g.:

@sample1:1
CGTACGTACGTACGTACGTACGTACGTACGTACGTACGTACGTACGTACGTACGTACGTA
+
IIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIII

6. Validate and preview#

oxo-flow validate qc-pipeline.oxoflow
# ✓ qc-pipeline.oxoflow — 4 rules, 3 dependencies

oxo-flow dry-run qc-pipeline.oxoflow
oxo-flow v0.23.2 — Rust-native bioinformatics pipeline engine
INFO Auto-discovered 2 samples from pattern 'raw_data/{sample}_R1.fastq.gz'
Plan: would run: 7 | skip: 0 | completed: 0 (DAG size: 7)
  1. fastp_trim_auto-discovered_sample1
     threads=4
     env=conda
     memory=8G
     outputs: ["results/trimmed/sample1_R1.fastq.gz", "results/trimmed/sample1_R2.fastq.gz", ...]
     command: mkdir -p results/trimmed
fastp --in1 raw_data/sample1_R1.fastq.gz --in2 raw_data/sample1_R2.fastq.gz --out1 results/trimmed/sample1_R1.fastq.gz --out2 results/trimmed/sample1_R2.fastq.gz --html results/trimmed/sample1_fastp.html --json results/trimmed/sample1_fastp.json --thread 4

  2. fastp_trim_auto-discovered_sample2
     threads=4
     env=conda
     memory=8G
     outputs: ["results/trimmed/sample2_R1.fastq.gz", "results/trimmed/sample2_R2.fastq.gz", ...]

  3. fastqc_raw_auto-discovered_sample1
     threads=4
     env=conda
     memory=8G
     outputs: ["results/fastqc/sample1_R1_fastqc.html", ...]

  4. fastqc_raw_auto-discovered_sample2
     threads=4
     env=conda
     memory=8G
     outputs: ["results/fastqc/sample2_R1_fastqc.html", ...]

  5. fastqc_trimmed_auto-discovered_sample1
     threads=4
     env=conda
     memory=8G
     outputs: ["results/fastqc_trimmed/sample1_R1_fastqc.html", ...]

  6. fastqc_trimmed_auto-discovered_sample2
     threads=4
     env=conda
     memory=8G
     outputs: ["results/fastqc_trimmed/sample2_R1_fastqc.html", ...]

  7. multiqc
     threads=1
     env=conda
     memory=8G
     outputs: ["results/multiqc/multiqc_report.html"]
     command: mkdir -p results/multiqc
multiqc results -o results/multiqc --force

Summary: 7 rules, total 25 threads declared, max 4 threads/rule
         7 rule(s) with memory requirements
         1 sample group(s), 0 pair(s)

To execute:  oxo-flow run qc-pipeline.oxoflow -j 2

The dry-run has expanded the {sample} wildcard into per-sample tasks: the three per-sample rules became one task per discovered sample (_auto-discovered_sample1, _auto-discovered_sample2), while multiqc — whose inputs and outputs contain no {sample} — stays a single task: 7 tasks in total. (The transcript above is abridged — real output also prints a checkpoint line, per-task input ✓/✗ status for concrete paths, and a command: line for every task.) Missing inputs do not fail dry-run — it exits 0; run is what fails.

validate warns; it does not gate

The ✓ qc-pipeline.oxoflow — 4 rules, 3 dependencies line only reports that the workflow parsed — validate exits 0 even when it warns. A missing input file, a sample_pattern that matches no files, or a {sample} wildcard with no sample source declared (no sample_pattern, [[sample_groups]], or [[pairs]]) — all appear as warnings under that line. Read them: run is what fails — it exits non-zero when a rule's input is absent or its wildcards cannot be bound.

Aggregation rules are expanded per sample too — and the duplicates race

multiqc became two tasks, but both aggregate the same results directory and write the same results/multiqc/ outputs. Because they sit at the same DAG level they run concurrently, and two MultiQC processes wiping/writing results/multiqc/multiqc_data at the same time fail with FileExistsError/FileNotFoundError — the tutorial workflow as written above fails on rerun. Never leave a {sample} in an aggregation rule's inputs when its outputs contain no per-sample path.

The fix is to key the aggregation to one sample's output and let MultiQC itself re-scan the whole results/ tree (it picks up every sample's reports):

[[rules]]
name = "multiqc"
input = [
    "{config.results_dir}/fastqc/sample1_R1_fastqc.html",
    "{config.results_dir}/fastqc_trimmed/sample1_R1_fastqc.html"
]
output = [ "{config.results_dir}/multiqc/multiqc_report.html" ]

The rule no longer contains {sample}, so it becomes a single task that waits for fastqc_raw/fastqc_trimmed through file matching, and the finished report covers all samples. (The tutorial workflow above has been updated to this form; when multiple tasks would write the same output path, the engine now flags it at plan time with a SCI-AGG-RACE preflight warning suggesting expand_inputs — see issue #443.)

The suggested -j 2 comes from dividing the machine's CPU threads by the workflow's maximum per-rule thread declaration (10 / 4 → 2, rounded down) — running more jobs than that would oversubscribe the CPU. If your rules are I/O-bound you can raise it.

Listing order reflects parallel levels

fastp_trim and fastqc_raw are listed adjacent because they are independent and will run in parallel at the same DAG level. Rules within a level are sorted alphabetically; fastqc_trimmed waits for fastp_trim, and multiqc waits for all three.


7. Visualize the DAG#

oxo-flow graph qc-pipeline.oxoflow

8. Run with parallel execution#

oxo-flow run qc-pipeline.oxoflow -j 2

The -j 2 flag allows up to 2 jobs to run concurrently — matching the suggestion from dry-run (machine threads ÷ per-rule threads). oxo-flow will execute the four independent level-0 tasks (two fastp_trim, two fastqc_raw) two at a time, then the two fastqc_trimmed tasks, then the single multiqc task (which re-scans the whole results/ tree, so its report covers both samples).

The QC numbers themselves are meaningless

The reads are synthetic 60 bp sequences, so FastQC/fastp run to completion but report nothing biologically useful — the tutorial's goal is a pipeline that executes end to end, not real QC. Point sample_pattern at real FASTQs (or drop them into raw_data/ and re-run) to get meaningful metrics.


Key Concepts Covered#

Concept Where you saw it
Workflow metadata [workflow] section with name, version, description
Configuration variables [config] section referenced as {config.samples_dir}
Defaults [defaults] section applied to all rules
Per-rule overrides multiqc rule overrides threads = 1, superseding [defaults] threads = 4
Environment specs environment = { conda = "envs/qc.yaml" }
Wildcard patterns {sample} in file paths
Sample auto-discovery sample_pattern = "raw_data/{sample}_R1.fastq.gz" in [workflow]
Multi-line shell Triple-quoted strings with """
Automatic dependencies Input/output matching across rules (all edges in this tutorial)
Explicit dependencies depends_on field — shown in the info box as an alternative for output-less rules

Next Steps#