Skip to content

oxo-flow dry-run#

Simulate execution without running any commands. Shows the execution plan, rule order, and expanded shell commands — and, when a checkpoint exists, predicts the actual incremental plan: which rules would re-run, which cascade downstream, and which stay protected.


Usage#

oxo-flow dry-run [OPTIONS] [WORKFLOW] [KEY=VALUE]...

plan is an alias for dry-run — oxo-flow plan [OPTIONS] [WORKFLOW] [KEY=VALUE]... is byte-identical, flag for flag. The documentation speaks of the execution plan (this page's "Plan:" headline, the checkpoint preview's plan array in --json); the subcommand now matches that vocabulary (issue #831).

dry-run accepts the same configuration inputs as run — --arg, trailing KEY=VALUE overrides, --profile, --rerun, and --resume-failed — and predicts the execution set with the exact same machinery (issue #77 parity contract). The preview also mirrors the executor's freshness gate: a rule with no checkpoint entry whose outputs are up to date is predicted as skipped, exactly as run would skip it.


Arguments#

Argument Description
[WORKFLOW] Path to the .oxoflow workflow file. Optional — if not specified, auto-discovery searches for: (1) main.oxoflow in current directory, (2) alphabetically first *.oxoflow file in current directory.
[KEY=VALUE]... Direct config overrides as trailing positionals — the same forms run accepts (KEY=VALUE, --KEY=VALUE, and the declared-key-only --KEY VALUE; see run). Command flags must come before the overrides.

Options#

Option Short Description
--target -t Run only specific target rules and their dependencies (repeatable, prefix matching)
--module — Run one include module and the producers of its declared inputs (repeatable; unions with --target). Module names are the include's name field or its file stem
--samples <LIST> — Sample selection: @path replaces the workflow's samples from a samplesheet, +@path appends (same-name groups merge, new groups added); names filter (or declare when the workflow ships no samples), first:N (pilot) and ready (samples whose entry inputs are complete) filter. Repeatable, comma-separated
--workdir <DIR> -d Resolve relative paths against this directory (default: the workflow file's directory)
--profile <NAME> — Execution profile loaded from profiles/<NAME>.toml — the SAME merge semantics as run
--arg <KEY=VALUE> — Set a workflow config value (overrides [config] defaults). Repeatable
KEY=VALUE… — Direct config overrides as trailing positionals (KEY=VALUE, --KEY=VALUE, --KEY VALUE) — the same forms run accepts; command flags must come before them
--rerun — Preview run --rerun: every rule in the execution set is forced (when-false rules still skip)
--resume-failed — Preview run --resume-failed: failed rules re-run, completed rules stay skipped
--skip-ref-build — Skip automatic reference/index building (assume pre-built) — the preview otherwise lists required builds
--cache-dir <PATH> — Directory for caching environment setup state — accepted for 1:1 run→dry-run transcription (the preview reads no env cache, so it is ignored). A misordered --cache-dir after overrides gets the same ordering hint as under run
--ai — Enable AI-powered analysis of the workflow
--ai-max-retries <N> — Maximum AI analysis attempts when a call fails (default: 1; only a failed call is retried)
--verbose -v Enable debug-level logging

Scientific Preflight#

Every dry-run runs deterministic, evidence-backed scientific checks on the workflow design and prints findings (e.g. SCI-VQSR-COHORT, SCI-MUTECT2-TUMOR-ONLY). With --ai, the findings are also passed to the model for a plain-language explanation.

The SCI-FEATURECOUNTS-STRAND check fires only when featureCounts is actually invoked (a command word) without a strandness flag — a rule that merely passes a .../featurecounts output path to another tool stays silent (issue #441). It also stays silent when the config declares strandedness = "unstranded" (the default -s 0 is then deliberate), and its advice is evidence-driven: determine strandedness with RSeQC infer_experiment.py or the samplesheet's strandedness column before setting -s 1/-s 2.

The SCI-AGG-RACE check flags an aggregation rule whose inputs/shell/when (or expand_inputs patterns) activate a fan-out dimension — {sample} / {group}, a pair wildcard, a [[values]] table name, or an output_pattern producer's fresh wildcard — while the declared outputs are not keyed by that dimension (issue #443; per-dimension keying since #829). Expansion creates one instance per fan-out element; every instance of an unkeyed dimension bakes the SAME concrete output path: concurrently they race, sequentially they duplicate work N−1 times. The check is per-dimension, so a rule whose outputs carry one dimension's wildcard (say {pair_id}) while another active dimension (say a [[values]] table referenced from expand_inputs) is unkeyed still fires — the warning names each unkeyed dimension. The remedy is either the documented aggregation idiom (bind the dimension through variables/expand_inputs so the rule becomes a single instance) or keying the outputs by the dimension's wildcard for a genuine per-element rule. The detector mirrors the engine's fan-out semantics, so fully keyed rules, input_groups rules, output_pattern producers, consumers keyed by the producer's fresh wildcard, and when-gated-off rules stay silent.

# A 2-sample pilot of a VQSR workflow fails for scientific reasons —
# the preflight says so before any compute is spent
oxo-flow dry-run pipeline.oxoflow --samples first:2

Sample Readiness#

Every dry-run on a sample-scoped workflow reports which samples have complete entry inputs and which are still waiting for data — designed for incremental data arrival, when a sequencing center delivers fastq files in batches:

$ oxo-flow dry-run pipeline.oxoflow
Plan: would run: 100 | skip: 0 | completed: 0 (DAG size: 100)
Sample readiness: 87/100 complete, 13 waiting
    ⏳ NA12891 (missing: data/NA12891_R2.fastq.gz)
    ⏳ NA12892 (missing: data/NA12892_R2.fastq.gz)
    … and 11 more waiting

Rules are judged per sample:

  • A sample is ready when every external input belonging to it exists. External inputs are rule inputs (after wildcard and {config.x} expansion) that the workflow itself does not produce; intermediate products are the DAG's job, so they are never checked.
  • optional = true rules do not block readiness — the executor skips them when their inputs are absent.
  • Missing files that belong to no specific sample (shared references) are reported as workflow-level inputs that block every sample.
  • Relative paths resolve against the workflow file's directory — the same place rules run from — so the report is accurate even when you invoke dry-run from another directory.
  • --samples ready previews only the ready samples, but the readiness section still covers the whole cohort so you can see what was left out.

With --json the same report is machine-readable:

"samples": {
  "total": 100, "ready": 87, "waiting_count": 13,
  "ready_names": ["NA12878", "…"],
  "waiting": [{"name": "NA12891", "missing": ["data/NA12891_R2.fastq.gz"]}],
  "missing_global": []
}

See run for the matching --samples ready execution mode.

Checkpoint-Aware Rerun Preview#

dry-run loads .oxo-flow/checkpoint.json read-only and classifies every rule in the execution set exactly the way run would — same config-impact fingerprints, same input manifests, same DAG downstream closure — so the preview matches what an actual run will do. Without a checkpoint the same classification still runs against an empty state (every rule "never completed", when conditions still honored). A warning makes the absent checkpoint explicit instead of implying a completed-run state:

⚠ no checkpoint at ./.oxo-flow/checkpoint.json — treating every rule as never completed
$ oxo-flow dry-run pipeline.oxoflow --samples NA12891
Plan: would run: 12 | skip: 0 | completed: 705 (DAG size: 717)
Checkpoint: ./.oxo-flow/checkpoint.json (modified 2026-08-10 14:32)
  completed: 705 | will run: 12 | will skip: 0 | protected (outside this run): 693
  rerun cascade: trim_cohort_NA12891 → align_cohort_NA12891 → combine_gvcfs → genotype_gvcfs → vqsr_snps
  1. trim_cohort_NA12891  [run: input changed]
  2. align_cohort_NA12891  [rerun: downstream of trim_cohort_NA12891]
  ...
  12. vqsr_snps  [rerun: downstream of trim_cohort_NA12891]

The other 99 samples' 693 completed rules are outside this execution set — their work stays untouched, counted as protected.

Per-rule status markers:

Marker Meaning
[run: never completed] No checkpoint entry — it will execute
[run: input changed] Input files differ from the manifest recorded at completion
[run: config changed] Config value or rule definition changed since completion
[run: outputs missing] Declared outputs no longer exist
[rerun: downstream of X] Was completed, but sits downstream of a rule that will execute (the cascade)
[rerun: upstream of X] A completed producer regenerates first because rule X needs its (tombstoned or missing) outputs — lazy cascade-up
[skip: up to date] Checkpoint hit — work stays protected
[skip: when condition false] The rule's when condition evaluates to false against the merged config — run skips it regardless of invalidation state

The headline (Plan: would run: N | skip: M | completed: K) answers the question that matters before any run: what happens next — how many rules execute, how many are skipped, and how much completed work the checkpoint holds. The total rule count in the DAG is shown as DAG size, not as the headline number (issue #432: DAG: 391 rules would execute read as "everything re-runs" while 381 of them were actually up to date — a costly misread during triage). The summary line also answers how much prior work survives (protected): rules outside this execution set stay untouched. The cascade line makes the infection chain visible — one sample's data change reaching the queue-level rules is exactly the part users cannot see from the DAG alone.

--profile <NAME> applies the SAME merge run uses (profile values fill in config keys the workflow does not set), so a preview computed with the profile matches what a profiled run would invalidate — and a preview without it flags exactly the drift. When references are declared and their build outputs are missing, the preview lists them ("References: N reference build(s) would run"); pass --skip-ref-build to assume they are pre-built, mirroring the run flag.

Temporary rules (temporary = true) are modeled exactly like run treats them: a tombstoned rule whose outputs were deleted by design shows [skip: up to date] while no dependent needs them, and flips to [rerun: upstream of X] the moment a dependent will execute again — regenerating the intermediate is part of the predicted plan, so the preview and the actual run stay identical. See run for the execution-side semantics.

The preview is strictly read-only and never mutates the checkpoint; it is orthogonal to run --rerun (which forces execution) — the preview only predicts, it changes no execution semantics. With --json the same prediction is machine-readable:

"checkpoint_preview": {
  "path": ".oxo-flow/checkpoint.json",
  "modified": "2026-08-10 14:32:00",
  "completed_total": 705,
  "summary": {"will_run": 12, "will_skip": 0, "protected_outside": 693},
  "plan": [
    {"name": "trim_cohort_NA12891", "status": "run-input-changed", "cascaded_from": null},
    {"name": "combine_gvcfs", "status": "run-cascaded", "cascaded_from": "trim_cohort_NA12891"}
  ],
  "cascade_chains": [["trim_cohort_NA12891", "combine_gvcfs", "genotype_gvcfs", "vqsr_snps"]]
}

Top-level fields: "profile" (the --profile name, when given) and "reference_builds" (reference names whose build outputs are missing — --skip-ref-build empties the list).

Pixi manifests declared by pending rules are preflighted exactly like run does at start-up (issue #848): every non-skip rule's pixi = "…" spec is resolved against the workflow directory and existence-checked. Missing manifests appear under the top-level key "env_manifest_issues" (an array of {"rule", "spec"} objects, spec already resolved), the human stderr output gains a matching block, and — like a cycle or a parse error — a non-empty list fails the command with a non-zero exit after the JSON document is printed, so machine consumers read the JSON from stdout and the failure reason from stderr or the exit code:

"env_manifest_issues": [
  {"rule": "trim_cohort_NA12891", "spec": "envs/trim/pixi.toml"}
]

Status values: run-never-completed, run-input-changed, run-config-changed, run-outputs-missing, run-cascaded, run-cascaded-upstream, run-forced, skip, skip-fresh, skip-when-condition.

--json also emits the execution-plan surface at the top level (schema_version: 1) — the structured mirror of the human stderr plan, for CI scripts and tooling that today grep the colored text:

"plan": [
  {"name": "trim_cohort_NA12891", "status": "run", "reason": "input changed",
   "cascaded_from": null, "threads": 4, "memory": "16G",
   "environment": "conda", "command": "trimmomatic PE …",
   "inputs": […], "outputs": […],
   "inputs_expanded": […], "outputs_expanded": […]}
],
"summary": {"would_execute": 12, "will_skip": 0, "total_rules": 12},
"suggested_jobs": 2,
"sample_groups": [{"name": "cohort", "samples": […]}],
"pairs": [{"pair_id": "…", "experiment": "…", "control": "…"}]

status ∈ run | skip | rerun matches the stderr bracket prefix; reason carries the status text. The human stderr output is unchanged by --json. inputs/outputs carry the declared patterns as written; inputs_expanded/outputs_expanded are the exact per-instance paths the engine will touch (with {config.x} resolved) — the same expansion the human listing shows.

Examples#

Preview with auto-discovery#

# Auto-discover workflow in current directory
oxo-flow dry-run

Preview a specific workflow#

oxo-flow dry-run pipeline.oxoflow

Preview a specific target rule and its dependencies#

oxo-flow dry-run pipeline.oxoflow -t align

Preview multiple target rules#

oxo-flow dry-run pipeline.oxoflow -t align -t sort_bam

With verbose output#

oxo-flow dry-run pipeline.oxoflow -v

Checkpoint re-entry in previews#

Recorded re-entries whose checkpoint rule is still up-to-date replay into the preview: the preview shows the same static plan a real run would execute (round-1 instances appear as up-to-date skips). Checkpoint rules that may add instances at runtime are listed under the reentry section of --json (recorded + possible). See Workflow Format.

Output#

⚠ no checkpoint at ./.oxo-flow/checkpoint.json — treating every rule as never completed
Checkpoint: ./.oxo-flow/checkpoint.json (modified unknown time)
Plan: would run: 3 | skip: 0 | completed: 0 (DAG size: 3)
  1. generate_data  [run: never completed]
     threads=1
     outputs: ["data/raw.csv"]
     command: mkdir -p data
echo 'id,name,value' > data/raw.csv
for i in $(seq 1 100); do echo "$i,item_$i,$((i * 37 % 1000))"; done >> data/raw.csv

  2. transform  [run: never completed]
     threads=1
     outputs: ["data/filtered.csv"]
     command: head -1 data/raw.csv > data/filtered.csv
awk -F',' 'NR>1 && $3 > 500' data/raw.csv >> data/filtered.csv

     input ✗: data/raw.csv
  3. summarize  [run: never completed]
     threads=1
     outputs: ["results/summary.txt"]
     command: mkdir -p results
total=$(tail -n +2 data/filtered.csv | wc -l)
echo 'Welcome to oxo-flow' > results/summary.txt
echo "Filtered records: $total" >> results/summary.txt
echo "Generated by oxo-flow file-pipeline" >> results/summary.txt

     input ✗: data/filtered.csv

Summary: 3 rules, total 3 threads declared, max 1 threads/rule

To execute:  oxo-flow run 02_file_pipeline.oxoflow -j 1

The human-readable plan goes to stderr (including the Summary: and To execute: lines) so it stays visible above redirected job output; with --json the machine-readable report goes to stdout instead. The sample above is a first run on fresh state: a ⚠ no checkpoint … warning and a Checkpoint: line precede the plan, and each rule carries a [run: never completed] marker (see below); with an existing checkpoint the warning is replaced by checkpoint status details and completed rules instead show [skipped]/re-entry markers. The sample-scoped output also includes a Sample readiness: section.


Notes#

  • The workflow file is optional; if not specified, auto-discovery searches for main.oxoflow first, then any *.oxoflow file alphabetically
  • If no .oxoflow file is found, an error message suggests running oxo-flow init to create one
  • No shell commands are executed — the dry-run is read-only
  • The checkpoint is loaded read-only too — dry-run never saves, baselines nothing, and invalidates nothing on disk
  • The preview mirrors run's incremental semantics; it is orthogonal to run --rerun (which forces execution) and run's config-change invalidation — see Run for those
  • Shell commands are shown in full — they may span multiple lines
  • The environment type (conda, docker, etc.) is shown for each rule
  • Thread and resource settings are displayed per rule
  • Use dry-run to verify your workflow before committing compute resources
  • When --target is specified, only the named rules and all rules they depend on (transitively) are shown — downstream rules are excluded