Writing Custom Scripts#
This tutorial bridges the gap between running standard shell commands and writing complex logic within oxo-flow. Often, a single shell command isn't enough to process bioinformatics data, and you'll need to embed Python, R, or Bash scripts into your workflow.
The script Directive#
Instead of a shell command, a rule can declare a script field that points to an external script file. The interpreter is auto-detected from the file extension, and the script runs inside the rule's declared environment.
A Complete Example#
Here is a runnable two-step pipeline where each step is a Python script. Every rule declares its input and output arrays — that is what connects script-based rules into the DAG.
1. The workflow (count.oxoflow)#
[workflow]
name = "count-reads"
version = "0.1.0"
[config]
sample = "SAMPLE_01"
[[rules]]
name = "count"
input = ["raw/{config.sample}.fastq"]
output = ["counts/{config.sample}.count.txt"]
script = "scripts/count_reads.py {input[0]} --min-quality {params.min_q} -o {output[0]}"
[rules.params]
min_q = 20 # ← defines {params.min_q} used above
[rules.environment]
conda = "envs/py.yaml"
[[rules]]
name = "summarize"
input = ["counts/{config.sample}.count.txt"] # ← consumes count's output: DAG edge
output = ["reports/{config.sample}.summary.txt"]
script = "scripts/summarize.py {input[0]} -o {output[0]}"
[rules.environment]
conda = "envs/py.yaml"
summarize depends on count automatically — its input matches count's output, so the DAG engine infers the edge. No explicit declaration needed.
2. The environment file#
Both rules declare environment = { conda = "envs/py.yaml" }, so create that file first — the run will fail with EnvironmentFileNotFound otherwise:
3. The scripts#
scripts/count_reads.py:
#!/usr/bin/env python3
"""Count reads above a minimum quality threshold."""
import argparse
parser = argparse.ArgumentParser()
parser.add_argument("fastq", help="input FASTQ file")
parser.add_argument("--min-quality", type=int, default=20)
parser.add_argument("-o", dest="out", required=True, help="output file")
args = parser.parse_args()
count = 0
with open(args.fastq) as f:
for line in f:
if line.startswith("@"):
count += 1
with open(args.out, "w") as f:
f.write(f"{count}\n")
print(f"counted {count} reads")
scripts/summarize.py:
#!/usr/bin/env python3
"""Turn a count file into a one-line summary."""
import argparse
parser = argparse.ArgumentParser()
parser.add_argument("count_file", help="input count file")
parser.add_argument("-o", dest="out", required=True, help="output file")
args = parser.parse_args()
with open(args.count_file) as f:
n = f.read().strip()
with open(args.out, "w") as f:
f.write(f"total reads: {n}\n")
print(f"wrote summary: {n} reads")
4. Run it#
Create a tiny FASTQ file to process, then run the pipeline:
mkdir -p raw
printf '@read1\nACGTACGT\n+\nIIIIIIII\n@read2\nTGCATGCA\n+\nIIIIIIII\n' > raw/SAMPLE_01.fastq
oxo-flow run count.oxoflow
# ✓ count (0.1s)
# ✓ summarize (0.1s)
cat reports/SAMPLE_01.summary.txt
How {params.*} Works#
{params.<key>} placeholders refer to the rule's [rules.params] table — key-value pairs scoped to a single rule (unlike {config.*}, which is global):
script = "analyze.py --min-quality {params.min_q} --genome {params.genome} {input[0]}"
# Expands to: analyze.py --min-quality 20 --genome hg38 raw/SAMPLE_01.fastq
| Placeholder | Scope | Defined in |
|---|---|---|
{params.<key>} |
One rule | [rules.params] |
{config.<key>} |
Whole workflow | [config] |
{input[N]} / {input} |
One rule | input array |
{output[N]} / {output} |
One rule | output array |
{threads} / {memory} |
One rule | threads / memory fields |
All placeholders are expanded before the script is launched, so the script itself never sees {...} syntax — it receives concrete values as command-line arguments.
Passing Files to Scripts#
Scripts receive file paths as ordinary command-line arguments. Use {input[N]} / {output[N]} placeholders in the script string to pass them:
For scripts that prefer stdin/stdout streams, wrap them in shell instead:
Interpreter Detection#
The interpreter is detected automatically, in this order:
- Explicit
interpreterfield on the rule - Custom
[workflow.interpreter_map]in the workflow metadata - Built-in defaults based on file extension:
| Extension | Interpreter |
|---|---|
.py / .py3 |
python / python3 |
.R / .r |
Rscript |
.sh / .bash |
bash |
.jl |
julia |
.pl |
perl |
.rb |
ruby |
.qmd / .Rmd |
quarto render |
.ipynb |
jupyter nbconvert --execute |
.smk |
snakemake |
.nextflow |
nextflow run |
.wdl |
miniwdl run |
- Shebang line (if the file is executable)
Combining shell and script#
When both are declared, shell runs first, then script:
[[rules]]
name = "qc_and_report"
shell = "mkdir -p results/"
script = "scripts/qc_report.R {input[0]} results/"
Outputs are verified after both complete.
Scripts vs. Shell Commands#
shell |
script |
|
|---|---|---|
| Best for | Short commands, one-liners, pipes | Multi-step logic, complex programs |
| Language | Any shell (bash, sh) | Any language with an interpreter |
| Interpreter | Shell itself | Auto-detected from extension, interpreter field, or shebang |
| Dependency tracking | Same | Same |
| Output verification | Same — declared outputs must exist after completion | Same |
See Also#
- Script Execution (workflow format) — full reference for the
scriptfield - Custom Interpreters (
interpreter_map) — mapping extensions to interpreters - Command Reference — running workflows