Skip to content

02 — File Pipeline#

A linear three-step pipeline that generates data, transforms it, and produces a summary. This demonstrates how oxo-flow resolves dependencies automatically from file paths.

Concepts Covered

  • Multi-rule workflows with automatic dependency resolution
  • Input/output chaining (one rule's output is the next rule's input)
  • Config variables accessible in shell commands via {config.*}
  • Multi-line shell commands

Workflow Definition#

# examples/gallery/02_file_pipeline.oxoflow
# 02 — File Processing Pipeline
# A linear three-step pipeline: generate → transform → summarize.
# Demonstrates: multi-rule dependencies, input/output chaining, config variables.

[workflow]
name = "file-pipeline"
version = "1.0.0"
description = "Linear file processing pipeline with three sequential steps"
author = "oxo-flow examples"

[config]
greeting = "Welcome to oxo-flow"

[[rules]]
name = "generate_data"
output = ["data/raw.csv"]
shell = """
mkdir -p data
echo 'id,name,value' > {output[0]}
for i in $(seq 1 100); do echo "$i,item_$i,$((i * 37 % 1000))"; done >> {output[0]}
"""

[[rules]]
name = "transform"
input = ["data/raw.csv"]
output = ["data/filtered.csv"]
shell = """
head -1 {input[0]} > {output[0]}
awk -F',' 'NR>1 && $3 > 500' {input[0]} >> {output[0]}
"""

[[rules]]
name = "summarize"
input = ["data/filtered.csv"]
output = ["results/summary.txt"]
shell = """
mkdir -p results
total=$(tail -n +2 {input[0]} | wc -l)
echo '{config.greeting}' > {output[0]}
echo "Filtered records: $total" >> {output[0]}
echo "Generated by oxo-flow file-pipeline" >> {output[0]}
"""

Key Concepts#

Automatic Dependency Resolution#

oxo-flow builds a DAG by matching rule outputs to rule inputs:

  1. generate_data produces data/raw.csv
  2. transform requires data/raw.csv → depends on generate_data
  3. summarize requires data/filtered.csv → depends on transform

You never need to declare dependencies manually — they are inferred from file paths.

How the inference engine builds and topologically sorts the DAG is covered in the DAG engine reference.

Config Variables#

The [config] section defines key-value pairs that are accessible in shell commands via {config.key}:

[config]
greeting = "Welcome to oxo-flow"

In the shell: echo '{config.greeting}' expands to echo 'Welcome to oxo-flow'.

{config.x} is one of the three substitution mechanisms; see Wildcards for the others.

Multi-Line Shell Commands#

Use triple-quoted strings ("""...""") for multi-line shell commands. Each line is executed as part of a single shell invocation.

Running the Workflow#

Validate#

$ oxo-flow validate examples/gallery/02_file_pipeline.oxoflow
✓ examples/gallery/02_file_pipeline.oxoflow — 3 rules, 2 dependencies

Dry-Run#

$ oxo-flow dry-run examples/gallery/02_file_pipeline.oxoflow
oxo-flow v0.23.2 — Rust-native bioinformatics pipeline engine
Plan: would run: 3 | skip: 0 | completed: 0 (DAG size: 3)
  1. generate_data
     threads=1
     outputs: ["data/raw.csv"]
     command: mkdir -p data
echo 'id,name,value' > data/raw.csv
for i in $(seq 1 100); do echo "$i,item_$i,$((i * 37 % 1000))"; done >> data/raw.csv

  2. transform
     threads=1
     outputs: ["data/filtered.csv"]
     command: head -1 data/raw.csv > data/filtered.csv
awk -F',' 'NR>1 && $3 > 500' data/raw.csv >> data/filtered.csv

     input ✗: data/raw.csv
  3. summarize
     threads=1
     outputs: ["results/summary.txt"]
     command: mkdir -p results
total=$(tail -n +2 data/filtered.csv | wc -l)
echo 'Welcome to oxo-flow' > results/summary.txt
echo "Filtered records: $total" >> results/summary.txt
echo "Generated by oxo-flow file-pipeline" >> results/summary.txt

     input ✗: data/filtered.csv

Summary: 3 rules, total 3 threads declared, max 1 threads/rule

To execute:  oxo-flow run examples/gallery/02_file_pipeline.oxoflow -j 1

Execute#

$ oxo-flow run examples/gallery/02_file_pipeline.oxoflow
  ✓ generate_data (0.0s)
  ✓ transform (0.0s)
  ✓ summarize (0.0s)

Done: 3 succeeded, 0 skipped, 0 failed
✓ 3 output files verified (2303B total)

DAG Visualization#

$ oxo-flow graph examples/gallery/02_file_pipeline.oxoflow
┌──────────────────────────────────────────────┐
│  Workflow DAG: 3 rules, 2 dependencies       │
│  Depth: 3, Width: 1, Critical path: 3 steps  │
└──────────────────────────────────────────────┘

Level 0 (sequential)
     generate_data
     │
     ▼
Level 1 (sequential)
     transform [depends: generate_data]
     │
     ▼
Level 2 (sequential)
     summarize [depends: transform]

Critical path: generate_data → transform → summarize

For embedding in docs/GitHub, the mermaid format renders the same DAG as this page shows:

$ oxo-flow graph examples/gallery/02_file_pipeline.oxoflow --format mermaid
graph LR
    n0["generate_data"]
    n1["transform"]
    n2["summarize"]
    n0 --> n1
    n1 --> n2
graph LR
    A[generate_data] --> B[transform]
    B --> C[summarize]

What's Next?#

Move on to Parallel Samples to learn how wildcards enable multi-sample parallel processing.