Skip to content

oxo-flow ai#

AI status, self-test, setup, and three-layer workflow explanation.

oxo-flow ai                 # quick status: provider, model, endpoint, sessions
oxo-flow ai test            # full self-test (provider round-trip)
oxo-flow ai setup           # interactive provider configuration wizard
oxo-flow ai explain wf.oxoflow

Options#

Option Short Description
--step <RULE> — Explain a single rule by name (with explain)
--level <LEVEL> — Explanation depth: beginner (default — jargon is defined) or expert (parameter-level, efficiency-focused)
--json — Machine-readable JSON output (with explain): the deterministic skeleton plus model-written prose fields

The three-layer explanation (ai explain)#

  1. Overview — what the workflow does, in plain language.
  2. Per-step detail — every rule, its role, and how steps connect.
  3. Scientific review — deterministic evidence-backed constraints (the same preflight rules dry-run prints) grounded to rule text, plus model-written prose. The grounding is deterministic — the model only writes prose around machine-derived facts.

Degraded mode#

ai explain never hard-fails on the model: when the provider is unreachable, errors, or is explicitly disabled, it emits the deterministic grounding skeleton (workflow identity, per-step order/description/tools/inputs/outputs/resources, and the knowledge-base grounding) with a note on stderr, and exits 0. The skeleton is the same data the model would have received — verification data, not hallucination — so scripts can rely on the JSON contract in every state.

  • OXO_FLOW_AI_PROVIDER=disabled explicitly disables the model and overrides any saved provider configuration: the skeleton is emitted with a "disabled" note and exit 0.
  • A failed provider call (dead endpoint, quota, timeout) degrades the same way, with the error explained in the note.

Embedded knowledge freshness#

oxo-flow ai reports how fresh the eight embedded knowledge sources are (Bioconda, bio.tools, commercial tools, EDAM ontology, nf-core modules, pipeline graph, skillgraph docs, bioSkills library) in a Knowledge freshness section: per-source record count, generation date, staleness in days, and whether the source is auto-updated (auto) or manually curated (manual):

Knowledge freshness:
  bioconda_tools (auto) 6487 records, generated 2026-09-01 (7 days)
  biotools_overlay (auto) 1925 records, generated 2026-08-23 (15 days)
  commercial_tools (auto) 25 records, generated 2026-09-01 (7 days)
  edam_terms (auto) 839 records, generated 2026-09-01 (7 days)
  nfcore_modules (auto) 2000 records, generated 2026-09-01 (7 days)
  pipeline_graph (auto) 543 records, generated 2026-09-01 (7 days)
  skillgraph_docs (auto) 77 records, generated 2026-09-01 (7 days)
  skills_index (auto) 562 records, generated 2026-09-01 (7 days)
  • Auto-updated sources older than 60 days are flagged STALE — the same threshold the release pipeline's staleness gate enforces, so a shipped binary never embeds auto-updated knowledge older than 60 days.
  • lookup_tool responses carry the same data date and record count as a freshness note, so agents can weigh how current the embedded database is.
  • See "Embedded Knowledge Freshness" in the AI CLI reference for the update cadence, the update GitHub Action, and the gate details.

The --ai flag on other commands#

One flag, a different action per command — the context is the command itself:

Command What --ai does
run --ai-recover AI error recovery: on rule failure, the model analyzes stderr and proposes a fix (note: this one is --ai-recover, not --ai)
dry-run --ai Plain-language analysis of the workflow (scientific preflight findings passed to the model)
validate --ai Semantic validation beyond structure — scientific plausibility checks
lint --ai Semantic linting on top of the deterministic best-practice rules
debug --ai Plain-language explanation of a rule's expanded shell command
report --ai AI result interpretation — a summary of execution outcomes, caveats, and next steps
template --ai Generate a workflow from a natural-language description (optionally grounded in --from-url/--from-file reference material)
env create --ai Generate a conda/pixi environment spec from a natural-language description

All AI features go through the configured provider (see oxo-flow ai); deterministic behavior never depends on the model — AI output is always additive prose or proposals, never silent engine decisions.

How template --ai generates a workflow#

template --ai runs the same generation agent as the web translate and chat surfaces — one shared persona in oxo-flow-ai, one orchestrator loop. The steps:

  1. Resolve the destination — -o is parsed before anything is sent to the provider, so a bad output path fails cheaply.
  2. Gather context — --from-url references are fetched through the SSRF-screened fetcher; --from-file files are read with bounded previews. Both enter the prompt as external reference material.
  3. Plan — the shared persona's system prompt encodes the .oxoflow schema (single-brace templates, [[rules]] arrays, depends_on, version-pinned conda packages) and explicitly bans the wrong dialects models tend to drift into (foreach, inputs = {...} maps, [resources] sections). Domain-matched bioSkills and any activated [ai] skills are injected as prompt sections.
  4. Ground with tools — the model queries the embedded knowledge bases on demand: lookup_tool for exact Bioconda names/versions, lookup_skill/lookup_pipeline for domain procedures and topologies. Non-read-only tools (e.g. MCP tools from activated tool skills) ask for interactive approval on stderr before running; non-interactive sessions refuse them.
  5. Validate against the engine — every draft is parsed by the core engine (WorkflowConfig). On failure the errors feed back to the model for correction; this loop is bounded by --ai-max-retries.
  6. Deliver — the validated TOML is written, the three static gates (validate / dry-run / lint) can then confirm it, and the AI session (token usage, tool calls) is archived to the ai_sessions/ directory of the resolved .oxo-flow data dir — the project-local .oxo-flow/ when one exists, else ~/.oxo-flow/.

Correction budget and failure behavior#

  • --ai-max-retries N sizes the orchestrator's combined tool + correction-round budget (default from [ai] config, 10 when unset — knowledge lookups and the correction pass share the budget, exploration- heavy models spend 7–10 rounds before drafting, and the loop stops as soon as a draft validates). When the first half of that budget burns on consecutive rounds whose assistant turns were pure tool calls (zero prose), the orchestrator nudges once: a warning rides the next tool result telling the model to consolidate what it has found and start producing the deliverable — measured on thinking-style backends whose draws otherwise spent every round on lookups and hit the cap with nothing written.
  • --ai-attempts N is the fresh-draw budget across generations (best-of-N with early exit): every attempt starts from a new draw, the deterministic gates decide, and later attempts are only paid when earlier ones fail. Default 1.
  • Before any model round is spent on a failed draft, mechanical error classes are repaired in code (e.g. a missing [config] declaration — E005); only what the fixer cannot repair reaches the model.
  • If the budget runs out after a pipeline was drafted, the command degrades instead of discarding: the generated TOML is extracted from the transcript and written with a prominent review warning — a paid-for artifact is never silently lost. Degraded delivery never writes structurally broken TOML. Extraction scans the transcript's toml-fenced drafts latest-first (correction rounds accumulate, and the last complete draft is the model's final word); a fenced candidate is accepted only when it parses as TOML and passes the structural floor (a [workflow] table plus every rule carrying a shell). A draft the model corrected by opening a fresh fence — without closing the stale one — cannot shadow the correction: the stray fence line closes the dangling draft and reopens. When no complete draft exists, a parseable-but-incomplete fenced fragment is kept as a best-effort fallback (still delivered under the review warning). With no closed fence at all (budget expired mid-block), the raw fallback from the first [workflow] line likewise only accepts a prefix that parses and passes the floor — an unterminated draft fails extraction, the artifact is not written, and the error explains why.
  • Provider failures (auth, quota, network) fail fast with the error. One carve-out: a transient transport failure (dropped connection, mid-transfer reset, or a response body that fails to decode — the classic flakes where one request out of a healthy session dies) is retried twice with a short backoff before surfacing, since retrying is what a user would do. Timeouts are not retried — they have their own knob (OXO_FLOW_AI_TIMEOUT_SECS) and a retry would deterministically re-timeout.

Scientist Team profile (opt-in)#

--ai-team-profile full (or [ai] team_profile = "full") wraps the same generation agent in three extra roles: a task contract that standardizes under-specified requests before generation, a deterministic Curator brief that injects Bioconda candidates, domain procedures and pipeline-graph transitions up front, and an independent review pass whose blocking findings trigger exactly one regeneration. Compact remains the default; the measured quality/cost trade-off lives in eval/frontier/ in the repository (see also AI CLI reference → Generation team profiles).

Generation tuning knobs#

Variable Default When to change
OXO_FLOW_AI_MAX_TOKENS 16384 Calibrated for thinking-style backends (e.g. DeepSeek or GLM behind an Anthropic-compatible endpoint): reasoning blocks spend against this ceiling, and the old 4096 default was consumed before any TOML was produced. Non-thinking models keep ≥3× headroom (measured: 1.9k–4.4k output tokens per generation). Older first-party claude-3-* models cap max_tokens at 4096–8192 server-side — lower the knob when targeting them on the real api.anthropic.com
OXO_FLOW_AI_TIMEOUT_SECS 300 Non-thinking generations complete in 12–20 s; a single thinking round routinely runs past two minutes (a measured GLM round took 145 s, so the old 120 s default killed mid-round) — raise further for slower endpoints, together with OXO_FLOW_AI_MAX_TOKENS

Both are consumed by the Anthropic Messages backend; the DeepSeek-native, OpenAI-compatible, and Ollama backends have their own budgets.

Quality evidence for these calibrations (gate-scored benchmark, before/after the generation unification) lives in eval/frontier/ in the repository. The architecture view of the shared harness — how the CLI, translate, and chat surfaces differ only in provider resolution, tool policy, and presentation — is in Architecture → AI Subsystem.