oxo-flow ai#
AI status, self-test, setup, and three-layer workflow explanation.
oxo-flow ai # quick status: provider, model, endpoint, sessions
oxo-flow ai test # full self-test (provider round-trip)
oxo-flow ai setup # interactive provider configuration wizard
oxo-flow ai explain wf.oxoflow
Options#
| Option | Short | Description |
|---|---|---|
--step <RULE> |
— | Explain a single rule by name (with explain) |
--level <LEVEL> |
— | Explanation depth: beginner (default — jargon is defined) or expert (parameter-level, efficiency-focused) |
--json |
— | Machine-readable JSON output (with explain): the deterministic skeleton plus model-written prose fields |
The three-layer explanation (ai explain)#
- Overview — what the workflow does, in plain language.
- Per-step detail — every rule, its role, and how steps connect.
- Scientific review — deterministic evidence-backed constraints
(the same preflight rules
dry-runprints) grounded to rule text, plus model-written prose. The grounding is deterministic — the model only writes prose around machine-derived facts.
Degraded mode#
ai explain never hard-fails on the model: when the provider is
unreachable, errors, or is explicitly disabled, it emits the
deterministic grounding skeleton (workflow identity, per-step
order/description/tools/inputs/outputs/resources, and the knowledge-base
grounding) with a note on stderr, and exits 0. The skeleton is the
same data the model would have received — verification data, not
hallucination — so scripts can rely on the JSON contract in every state.
OXO_FLOW_AI_PROVIDER=disabledexplicitly disables the model and overrides any saved provider configuration: the skeleton is emitted with a "disabled" note and exit 0.- A failed provider call (dead endpoint, quota, timeout) degrades the same way, with the error explained in the note.
Embedded knowledge freshness#
oxo-flow ai reports how fresh the eight embedded knowledge sources are
(Bioconda, bio.tools, commercial tools, EDAM ontology, nf-core modules,
pipeline graph, skillgraph docs, bioSkills library) in a Knowledge
freshness section: per-source record count, generation date, staleness in
days, and whether the source is auto-updated (auto) or manually curated
(manual):
Knowledge freshness:
bioconda_tools (auto) 6487 records, generated 2026-09-01 (7 days)
biotools_overlay (auto) 1925 records, generated 2026-08-23 (15 days)
commercial_tools (auto) 25 records, generated 2026-09-01 (7 days)
edam_terms (auto) 839 records, generated 2026-09-01 (7 days)
nfcore_modules (auto) 2000 records, generated 2026-09-01 (7 days)
pipeline_graph (auto) 543 records, generated 2026-09-01 (7 days)
skillgraph_docs (auto) 77 records, generated 2026-09-01 (7 days)
skills_index (auto) 562 records, generated 2026-09-01 (7 days)
- Auto-updated sources older than 60 days are flagged
STALE— the same threshold the release pipeline's staleness gate enforces, so a shipped binary never embeds auto-updated knowledge older than 60 days. lookup_toolresponses carry the same data date and record count as a freshness note, so agents can weigh how current the embedded database is.- See "Embedded Knowledge Freshness" in the AI CLI reference for the update cadence, the update GitHub Action, and the gate details.
The --ai flag on other commands#
One flag, a different action per command — the context is the command itself:
| Command | What --ai does |
|---|---|
run --ai-recover |
AI error recovery: on rule failure, the model analyzes stderr and proposes a fix (note: this one is --ai-recover, not --ai) |
dry-run --ai |
Plain-language analysis of the workflow (scientific preflight findings passed to the model) |
validate --ai |
Semantic validation beyond structure — scientific plausibility checks |
lint --ai |
Semantic linting on top of the deterministic best-practice rules |
debug --ai |
Plain-language explanation of a rule's expanded shell command |
report --ai |
AI result interpretation — a summary of execution outcomes, caveats, and next steps |
template --ai |
Generate a workflow from a natural-language description (optionally grounded in --from-url/--from-file reference material) |
env create --ai |
Generate a conda/pixi environment spec from a natural-language description |
All AI features go through the configured provider (see oxo-flow ai);
deterministic behavior never depends on the model — AI output is
always additive prose or proposals, never silent engine decisions.
How template --ai generates a workflow#
template --ai runs the same generation agent as the web translate and
chat surfaces — one shared persona in oxo-flow-ai, one orchestrator
loop. The steps:
- Resolve the destination —
-ois parsed before anything is sent to the provider, so a bad output path fails cheaply. - Gather context —
--from-urlreferences are fetched through the SSRF-screened fetcher;--from-filefiles are read with bounded previews. Both enter the prompt as external reference material. - Plan — the shared persona's system prompt encodes the
.oxoflowschema (single-brace templates,[[rules]]arrays,depends_on, version-pinned conda packages) and explicitly bans the wrong dialects models tend to drift into (foreach,inputs = {...}maps,[resources]sections). Domain-matched bioSkills and any activated[ai] skillsare injected as prompt sections. - Ground with tools — the model queries the embedded knowledge
bases on demand:
lookup_toolfor exact Bioconda names/versions,lookup_skill/lookup_pipelinefor domain procedures and topologies. Non-read-only tools (e.g. MCP tools from activated tool skills) ask for interactive approval on stderr before running; non-interactive sessions refuse them. - Validate against the engine — every draft is parsed by the core
engine (
WorkflowConfig). On failure the errors feed back to the model for correction; this loop is bounded by--ai-max-retries. - Deliver — the validated TOML is written, the three static gates
(
validate/dry-run/lint) can then confirm it, and the AI session (token usage, tool calls) is archived to theai_sessions/directory of the resolved.oxo-flowdata dir — the project-local.oxo-flow/when one exists, else~/.oxo-flow/.
Correction budget and failure behavior#
--ai-max-retries Nsizes the orchestrator's combined tool + correction-round budget (default from[ai]config, 10 when unset — knowledge lookups and the correction pass share the budget, exploration- heavy models spend 7–10 rounds before drafting, and the loop stops as soon as a draft validates). When the first half of that budget burns on consecutive rounds whose assistant turns were pure tool calls (zero prose), the orchestrator nudges once: a warning rides the next tool result telling the model to consolidate what it has found and start producing the deliverable — measured on thinking-style backends whose draws otherwise spent every round on lookups and hit the cap with nothing written.--ai-attempts Nis the fresh-draw budget across generations (best-of-N with early exit): every attempt starts from a new draw, the deterministic gates decide, and later attempts are only paid when earlier ones fail. Default 1.- Before any model round is spent on a failed draft, mechanical error
classes are repaired in code (e.g. a missing
[config]declaration — E005); only what the fixer cannot repair reaches the model. - If the budget runs out after a pipeline was drafted, the command
degrades instead of discarding: the generated TOML is extracted
from the transcript and written with a prominent review warning —
a paid-for artifact is never silently lost. Degraded delivery never
writes structurally broken TOML. Extraction scans the transcript's
toml-fenced drafts latest-first (correction rounds accumulate, and the
last complete draft is the model's final word); a fenced candidate is
accepted only when it parses as TOML and passes the structural floor
(a
[workflow]table plus every rule carrying ashell). A draft the model corrected by opening a fresh fence — without closing the stale one — cannot shadow the correction: the stray fence line closes the dangling draft and reopens. When no complete draft exists, a parseable-but-incomplete fenced fragment is kept as a best-effort fallback (still delivered under the review warning). With no closed fence at all (budget expired mid-block), the raw fallback from the first[workflow]line likewise only accepts a prefix that parses and passes the floor — an unterminated draft fails extraction, the artifact is not written, and the error explains why. - Provider failures (auth, quota, network) fail fast with the error.
One carve-out: a transient transport failure (dropped connection,
mid-transfer reset, or a response body that fails to decode — the
classic flakes where one request out of a healthy session dies) is
retried twice with a short backoff before surfacing, since retrying
is what a user would do. Timeouts are not retried — they have their
own knob (
OXO_FLOW_AI_TIMEOUT_SECS) and a retry would deterministically re-timeout.
Scientist Team profile (opt-in)#
--ai-team-profile full (or [ai] team_profile = "full") wraps the same
generation agent in three extra roles: a task contract that
standardizes under-specified requests before generation, a deterministic
Curator brief that injects Bioconda candidates, domain procedures and
pipeline-graph transitions up front, and an independent review pass
whose blocking findings trigger exactly one regeneration. Compact remains
the default; the measured quality/cost trade-off lives in eval/frontier/
in the repository (see also
AI CLI reference → Generation team profiles).
Generation tuning knobs#
| Variable | Default | When to change |
|---|---|---|
OXO_FLOW_AI_MAX_TOKENS |
16384 |
Calibrated for thinking-style backends (e.g. DeepSeek or GLM behind an Anthropic-compatible endpoint): reasoning blocks spend against this ceiling, and the old 4096 default was consumed before any TOML was produced. Non-thinking models keep ≥3× headroom (measured: 1.9k–4.4k output tokens per generation). Older first-party claude-3-* models cap max_tokens at 4096–8192 server-side — lower the knob when targeting them on the real api.anthropic.com |
OXO_FLOW_AI_TIMEOUT_SECS |
300 |
Non-thinking generations complete in 12–20 s; a single thinking round routinely runs past two minutes (a measured GLM round took 145 s, so the old 120 s default killed mid-round) — raise further for slower endpoints, together with OXO_FLOW_AI_MAX_TOKENS |
Both are consumed by the Anthropic Messages backend; the DeepSeek-native, OpenAI-compatible, and Ollama backends have their own budgets.
Quality evidence for these calibrations (gate-scored benchmark,
before/after the generation unification) lives in eval/frontier/ in
the repository. The architecture view of the shared harness — how the
CLI, translate, and chat surfaces differ only in provider resolution,
tool policy, and presentation — is in
Architecture → AI Subsystem.