Skip to content

oxo-flow resume#

Resume an interrupted workflow from a checkpoint file.

Usage#

oxo-flow resume [OPTIONS] <CHECKPOINT>

Description#

The resume command loads a checkpoint generated by a previous oxo-flow run and continues execution from where it left off. Rules that already completed successfully are skipped; only incomplete or previously-failed rules are re-executed.

The checkpoint file (.oxo-flow/checkpoint.json) is automatically created by oxo-flow run and stores the completion status and benchmarks of each rule, plus a config snapshot and per-rule fingerprints for change detection.

The checkpoint also records the working directory the original run executed in, so resume re-runs from the same place even when invoked from another directory — completed rules' outputs resolve identically and stay skipped. The recorded workflow path and workdir are stored as absolute paths; a legacy checkpoint without a recorded workdir falls back to the workflow file's directory.

Resuming goes through the same config-change impact analysis as run: if the workflow config or a rule definition changed since the checkpoint was written, only the affected rules and their downstream re-execute. See Config changes and precise invalidation. Input changes are detected the same way — glob, directory, and plain-file inputs are compared against the recorded manifests, and changed inputs re-execute the affected rules (see Input changes and manifest invalidation).

resume passes no CLI config overrides, so the effective config is the workflow's defaults. When the original run carried overrides (its recorded config snapshot differs), the resumed run prints a Config drift warning naming every key that changed and its old → new values — plan-time gated instances (pair/sample when, {meta.*} baking) re-judge on the defaults and may drift from the recorded instance set. Re-pass the original overrides with oxo-flow run (which auto-resumes) to reproduce the recorded instance set exactly.

Options#

Option Short Description
-j, --jobs <JOBS> — Number of parallel jobs (default: 1)
--ai-recover — Enable AI error recovery on rule failure
--ai-max-retries <N> — Maximum AI attempts for recovery when a call fails (default: 1; only a failed call is retried)
--workdir <DIR> -d Working directory to resume in (default: the one recorded in the checkpoint)
--keep-going -k Continue execution when a job fails (same semantics as run)
--timeout <SECS> — Timeout per job in seconds (0 = disabled), or a duration like 1h/30m
--no-report-snapshot — Skip the automatic report snapshot after the resumed run (same flag as run)
--background — Detach the resumed run into a background process and exit 0 immediately (see Background runs (--background))

Examples#

# Resume from the default checkpoint
oxo-flow resume .oxo-flow/checkpoint.json

# Resume with 4 parallel jobs
oxo-flow resume .oxo-flow/checkpoint.json -j 4

Notes#

  • The checkpoint stores a reference to the original workflow file path. If the workflow was moved or renamed, resume will fail with a clear error.
  • The recorded working directory is shown as Workdir: before execution. --workdir overrides it (e.g. when the project directory was moved).
  • oxo-flow run automatically resumes from the checkpoint if one exists. The standalone resume command is useful for explicitly re-running after inspecting the checkpoint state.
  • Use oxo-flow status to inspect the checkpoint before resuming.
  • Rules whose when gate recorded false are reported as skipped, not failed, in the resume banner (issue #690) — their stale failure entries from older checkpoints don't count toward the failed total, and the gate is re-judged on this resume anyway.
  • --background detaches the resumed run (see run's Background runs section).

See Also#