Skip to content

Web API#

oxo-flow includes a built-in REST API server for building, validating, running, and monitoring bioinformatics workflows. The server is built with axum and follows a domain-driven modular monolith architecture.


API Design Conventions#

  • Envelope: success responses are bare JSON objects/arrays; errors are { code, message, detail?, suggestion? }
  • Errors: { code: "E001", message, detail?, suggestion? }
  • Lists: GET /api/runs returns a cursor-paginated envelope { items, next_cursor, total } (limit ≤ 500, status/q filters); other list endpoints return bare arrays (≤ 100 items)
  • Versioning: /api/ prefix for all endpoints
  • Authentication: in team/hpc mode, protected endpoints accept an Authorization: Bearer <token> session token or an X-API-Key header. The generated OpenAPI spec at GET /api/openapi.json declares both bearerAuth and apiKey security schemes; public endpoints (health, login, license, openapi.json, etc.) require no auth.
  • Self-discoverable: OpenAPI 3.1 spec at GET /api/openapi.json. The spec is code-generated via utoipa from the #[utoipa::path] annotations on every route handler — there is no hand-maintained static file. crates/oxo-flow-web/tests/openapi_gate.rs is the drift gate: it asserts every route in the router appears in the generated spec, so a new route without an annotation fails CI.

Structured Error Format#

All errors return a unified JSON format:

{
  "code": "AUTH_REQUIRED",
  "message": "Authentication is required for this endpoint",
  "detail": "The request did not include a valid session token or Bearer token",
  "suggestion": "Please login at POST /api/auth/login to obtain a session token"
}

Rate limiting follows the same contract: over-limit requests get 429 {"code":"RATE_LIMITED", "message":"Rate limit exceeded", "detail":"retry in Ns", ...} with a Retry-After header (sliding window, 100 requests / 60 s per client IP by default).


Starting the Server#

# Mode 1: Personal (default) — SQLite, no auth, localhost
oxo-flow serve

# Mode 2: Team — auth enabled, network-facing
oxo-flow serve --mode team

# Mode 3: HPC — cluster-aware
oxo-flow serve --mode hpc

# Or via the standalone binary:
oxo-flow-web --mode personal -p 3000

System & Monitoring#

Health Check#

GET /api/health
Returns status, version, mode, uptime, component health (database, filesystem, scheduler, AI provider), resource usage, and license info.

components.ai_key_storage is "encrypted" when OXO_FLOW_MASTER_KEY is set, "plaintext" otherwise (issue #205).

components.engine (issue #579) reports the engine CLI the server spawns for runs, resolved once at startup by running <binary> --version:

{
  "status": "ok",
  "components": {
    "engine": {
      "version": "0.23.2",
      "path": "/usr/local/bin/oxo-flow",
      "version_mismatch": false
    }
  }
}

version_mismatch is true when the engine's major.minor differs from the server's — runs will execute under the engine's semantics, and the server log carries a drift warning with a hint to pin OXO_FLOW_BIN. The field is null when the probe failed (binary missing or unparseable --version output); the run log header remains the version record in that case.

System Info#

GET /api/system
Returns OS, architecture, PID, uptime, and version. Team/hpc modes require authentication (the endpoint left the anonymous whitelist in the v0.11 hardening).

The response also carries engine — the same engine-CLI info as /api/health's components.engine (issue #579).

Runtime Metrics#

GET /api/metrics
Returns real-time resource metrics: CPU%, memory (used/total/swap), active workflows, total requests, CPU count. Team/hpc modes require authentication.

Audit Logs#

GET /api/audit?days=7&page=1&per_page=50
Returns paginated structured audit entries:
{
  "entries": [{ "timestamp", "user", "action", "resource", "result" }],
  "days": 7,
  "page": 1,
  "per_page": 50,
  "total": 128
}
page defaults to 1 and per_page defaults to 50 (max 200). Team/hpc modes: admin-only (the trail spans every user's actions; personal mode keeps the localhost trust model).

Server-Sent Events#

GET /api/events
Accept: text/event-stream
SSE stream for real-time workflow execution events: terminal events (run_completed, run_failed, run_cancelled) plus per-rule events (rule_started, rule_completed, rule_failed, rule_skipped — parsed live from the engine's execution log). Keepalive is axum's 15-second comment ping; if the client falls behind the broadcast buffer, the stream carries a synthetic {"type":"lagged","data":{"missed":N}} event — refetch run state when you see it.

Team/hpc modes require a one-time ?ticket= minted at POST /api/events/ticket under normal bearer auth (EventSource cannot set an Authorization header; the ticket is single-use, expires in 30 s, and replaces the session token that used to travel — and leak — in the URL,

522). The stream is filtered to the subscriber's own#

runs — admins see everything. Events carry a user field (the owning user id, or null for system-wide events).

run_completed carries a summary field: the CLI's invalidation summary extracted from the execution log — config changes, edited rule definitions, and input-set changes that invalidated checkpoint records this run (null when the run had no invalidation activity).

The stream ends when the server shuts down (SIGTERM/Ctrl+C): the response is closed as part of the graceful drain, so an EventSource sees its error/close at shutdown rather than a hung connection (#572). Clients should reconnect with backoff after an unexpected close.


Authentication & Authorization#

Login#

POST /api/auth/login
Content-Type: application/json

{"username": "admin", "password": "admin"}
Returns session token, username, and role.

Credentials are environment-driven, not hardcoded. The example above only works when OXO_FLOW_DEV_MODE=1 (accepts password == username) or when matching env passwords are set: OXO_FLOW_ADMIN_PASSWORD, OXO_FLOW_USER_PASSWORD, and OXO_FLOW_VIEWER_PASSWORD define the admin/user/viewer accounts. Without one of these the login returns 401 — set the env vars when launching oxo-flow serve (see Web System Architecture).

Identity is pinned server-side (#516): the shared user/viewer passwords never mint or adopt the admin identity (or any existing API-created account) — admin signs in only with OXO_FLOW_ADMIN_PASSWORD, and a username that belongs to a managed account requires that account's own password. OAuth identities are namespaced (<provider>:<name>) and can never collide with local usernames. Sessions store the canonical users.id resolved at login; role lookups match by id only and a session whose user row is gone (deleted user, stale pre-upgrade session) is rejected with 401.

Check Session#

GET /api/auth/me
Authorization: Bearer <token>
Returns {"authenticated": true, "username": "admin", "role": "admin"} or {"authenticated": false}.

Sign Out#

POST /api/auth/logout
Authorization: Bearer <token>
Revokes the presented session server-side (#520) — the token stops authenticating immediately instead of living out its 24-hour TTL. The client should drop its local copy (logged_out: true on success). Deleting a user (DELETE /api/users/{id}) cascades: the deleted user's sessions and API keys are removed with the account.

License Status#

GET /api/license
Returns license type, validity, commercial use flag, and contact info.

Upload License#

POST /api/license/upload
Admin-only outside personal mode; the payload is bounded (64 KiB). The raw blob is not stored — only its size and SHA-256 land in the audit trail (the endpoint does not yet activate licenses) (#519).

Users (admin)#

GET    /api/users          # List users (admin only)
POST   /api/users          # Create user
DELETE /api/users/{id}     # Delete user

Pipeline Lifecycle (v0.8 /api/pipelines/*)#

The pre-v0.8 /api/workflows/* endpoints no longer exist in the source — the running server exposes only the /api/pipelines/* API below.

Parse#

POST /api/pipelines/parse
Content-Type: application/json

{"toml_content": "<workflow TOML>", "format_version": "1.0"}
Returns structured pipeline: pipeline_id, name, version, rules (with summaries), dag (nodes + edges), stats. Pure function, zero side effects.

Validate#

POST /api/pipelines/validate
Content-Type: application/json

{"pipeline_id": "...", "toml_content": "<TOML>", "base_dir": "inputs"}
Returns { valid, errors: [{ code, message, rule, suggestion }] }. base_dir (optional, for missing-input checks) must be relative (no ..) and resolves inside the acting user's workspace (#521).

Prepare#

POST /api/pipelines/prepare
Content-Type: application/json

{"toml_content": "<TOML>", "resolve_wildcards": true, "apply_defaults": true}
Expands wildcards, resolves environments. Returns expanded_rules_count, wildcard_combinations, environment_setup_cmds.

Build DAG#

POST /api/pipelines/dag
Content-Type: application/json

{"pipeline_id": "...", "toml_content": "<TOML>"}
Returns { nodes, edges, parallel_groups, critical_path, metrics } as structured JSON.

Format#

POST /api/pipelines/format
Content-Type: application/json

{"toml_content": "<TOML>"}
Returns canonical TOML formatting.

Lint#

POST /api/pipelines/lint
Content-Type: application/json

{"toml_content": "<TOML>"}
Returns { valid, errors: [{ code, message, rule, suggestion }] } — validation errors plus lint-level findings (e.g. missing description).

Stats#

POST /api/pipelines/stats
Content-Type: application/json

{"toml_content": "<TOML>"}
Returns aggregate pipeline statistics.

Diff#

POST /api/pipelines/diff
Content-Type: application/json

{"toml_a": "<TOML A>", "toml_b": "<TOML B>"}
Returns structured diffs: { diffs: [{ path, category, description, severity }] }.

Export#

POST /api/pipelines/export
Content-Type: application/json

{"toml_content": "<TOML>", "format": "docker|singularity"}
Generates Dockerfile or Singularity definition.

List / Save / Get / Update / Delete#

GET    /api/pipelines              # List pipelines (most recent first, up to 100)
POST   /api/pipelines              # Save new pipeline
GET    /api/pipelines/{id}         # Get pipeline with TOML content
PUT    /api/pipelines/{id}         # Update pipeline
DELETE /api/pipelines/{id}         # Delete pipeline
POST   /api/pipelines/search       # Search by name, tags, content

Execution & Runs#

Create Run#

PostgreSQL deployments: every /api/runs* endpoint returns 503 {"code": "RUNS_REQUIRE_SQLITE", "message", "detail", "suggestion"} — run execution requires a SQLite-backed server. Library/AI/auth domains remain available.

POST /api/runs
Content-Type: application/json

{"toml_content": "<workflow TOML>", "max_jobs": 4, "dry_run": false, "keep_going": false, "pipeline_id": "<uuid>", "cluster_id": "<cluster-id>", "samples": ["S1", "S2"], "targets": ["rule_name"]}
toml_content is the workflow source (required — the run is created from it, not from a saved pipeline); max_jobs, dry_run, keep_going, pipeline_id, cluster_id, samples, and targets are top-level fields (not nested under a config object). Run-mutating endpoints (create/cancel/pause/resume/retry/clean/resume-checkpoint) require the user role or above — the viewer role is read-only — and a pipeline_id must be readable by the acting user (#519). This flat shape is the typed CreateRunRequest schema — the OpenAPI spec declares it as the endpoint's requestBody, and the handler parses exactly that type, so the spec, the server, and the TypeScript client stay in lockstep. Returns { run_id, status: "queued", estimated_resources, execution_plan }.

Before spawning the executor the run stages the acting user's uploaded files (POST /api/files → workspace/users/<user>/inputs/) into the run's working directory, so metadata_file and relative input paths resolve the same way they do on the CLI. Files already present in the workdir are never overwritten.

pipeline_id (optional) only associates the run with a saved pipeline — the TOML still comes from toml_content. It makes the run execute in the pipeline's persistent working directory (workspace/users/<user>/pipelines/<id>), so the checkpoint survives across re-runs. Re-running with a changed config rebuilds exactly the rules referencing the changed keys (plus their DAG downstream) — the rest keep their checkpoint records and are skipped. Runs without pipeline_id get a fresh per-run sandbox and execute everything. Malformed or unknown pipeline ids are rejected (400 INVALID_PIPELINE_ID / 404 PIPELINE_NOT_FOUND).

Execution flags are forwarded to the CLI executor:

  • dry_run: true spawns the preview subcommand (oxo-flow dry-run) — nothing executes; the log shows the would-be plan.
  • max_jobs defaults to 4 at this endpoint (the CLI flag itself defaults to 1): an omitted max_jobs still spawns the CLI with -j 4, which is also the value the resource estimate assumes.
  • keep_going: true maps to --keep-going (the CLI also accepts -k).
  • samples (array of strings, optional) maps to --samples (comma-joined).
  • targets (array of strings, optional) maps to -t per entry.
  • cluster_id (optional) names a configured SSH cluster connection: the workdir is staged to that host and executed there (see Cluster Connections).

List Runs#

GET /api/runs?limit=100&cursor=<created_at>&status=<status>&q=<search>

Cursor pagination, not page numbers: pass the created_at of the last row of the previous page as cursor (created_at < cursor). limit defaults to 100 and is capped at 500. The response is the envelope { items, next_cursor, total }, where next_cursor: null means the last page.

Run Status#

GET /api/runs/{id}/status
Real-time status: { status, phase, nodes: [{ rule, status, started_at, duration_ms, exit_code }], timeline, resources }. A rule that did not run — when-gated off (when_verdicts == false) or abort-cancelled (rule_runs status cancelled, which records no set membership) — is surfaced as skipped, never as a forever-pending node (issues #739, #767).

DAG Status#

GET /api/runs/{id}/dag-status
DAG JSON with per-node live status. Color-coded: green=completed, blue=running, red=failed, gray=skipped. metrics.pending_nodes and the ETA count only rules that can still run — skipped rules (when-gated off or abort-cancelled) are skipped, not pending (issues #739, #767).

Diagnostics#

GET /api/runs/{id}/diagnostics
Deterministic error analysis: { failed_nodes: [{ rule, error_pattern, likely_cause, suggestions, auto_fixable, fix_action, relevant_log_lines }], warnings, resource_bottlenecks }. Uses 30+ deterministic error patterns — zero AI in this endpoint. Diagnosis reads only the newest 256 KiB of execution.log (seek-based tail, torn first line dropped), never the whole file (issue #710).

Run Detail#

GET /api/runs/{id}
The full run row: { id, user_id, pipeline_id, pipeline_snapshot, status, phase, workdir, started_at, finished_at, created_at } — the identity other run endpoints key off. (pipeline_id/pipeline_snapshot may be null for legacy rows; no pid — the host process id stays engine-internal.)

Run Preview#

GET /api/runs/{id}/preview
The instance-level dry-run plan (checkpoint_preview + execution order), persisted by the executor as dry-run-preview.json when a dry-run completes. Returns 404 NO_PREVIEW for runs that never produced one — or when the workdir path is not a regular file (issue #735: a rule-writable workdir may contain special files). Files above the 16 MiB server-written maximum return 413 PREVIEW_TOO_LARGE.

AI Status#

GET /api/runs/{id}/ai-status
Node-level execution statuses pulled from the run's checkpoint — the input the AI endpoints (explain/interpret) consume.

Run Report#

GET /api/runs/{id}/report
The deterministic report object the report page renders: run identity, node statuses with timings, output file tree. Zero AI — the same object answers report/ask and feeds report/visualize. Like the diagnostics endpoint, the report summarizes only the newest 256 KiB of execution.log (issue #710); the full log stays on GET /api/runs/{id}/logs. The report is failure-aware (issue #759): a failed run's narrative headline names the failed rule(s) and exit codes, and carries run_status + failed_rules: [{rule, exit_code, stderr_tail, stdout_tail}] from the checkpoint's structured failure data (stdout captured since #691, surfaced since #765).

Report Q&A#

POST /api/runs/{id}/report/ask
Content-Type: application/json

{"question": "which rules failed and why"}
Answers from the deterministic report (pattern matching over failed nodes, durations, file tree) — not an LLM call. Returns the answer string. On a failed run, EVERY question answers from the failure data (rule, exit code, bounded stderr excerpt — falling back to the stdout tail when stderr is empty) — never a completion claim (issues #759, #765).

Run Status exit codes and failure tails#

GET /api/runs/{id}/status, /dag-status, and /instances populate exit_code from the checkpoint's rule_runs records (issue #758): completed rules report 0, failed rules their recorded code — plus bounded stderr_tail and stdout_tail on failed nodes (issue #765; some tools print their root cause on stdout). The AI chat get_run_status tool carries the same fields.

Report Visualization#

POST /api/runs/{id}/report/visualize
Content-Type: application/json

{"type": "files" | "durations" | "volcano"}
Chart data built from real sources: files maps the run's file tree ({name, size_bytes}), everything else maps per-rule durations from the checkpoint. No canned rows.

Smart Retry#

POST /api/runs/{id}/retry
Content-Type: application/json

{"skip_succeeded": true}
The retry spawns a real run, so it pays the same #213 runs-per-minute limiter and the quota pre-flight as run creation — a retry that exceeds the budget answers 429 RUN_RATE_LIMITED / QUOTA_EXCEEDED (#519). from_rule is not supported: the engine has no "re-run from rule X downstream" mode (--target selects the upstream closure), so the field is rejected with 400 UNSUPPORTED_FIELD instead of being silently ignored.

The retry really executes: the returned new_run_id is a real run in the database (same workdir, same owner), spawned with --resume-failed --rerun so the failed rules re-execute despite their existing outputs and the checkpoint's cascade invalidation re-runs their downstream dependents. Returns { new_run_id, will_rerun: [...], will_skip: [...] } — plus an optional note when the plan is empty on a previously failed run (it died before executing any rule, e.g. a configuration error at spawn; the retry re-fails identically unless the error is fixed — issue #760).

Cancel#

POST /api/runs/{id}/cancel
Cancels a running/pending run.

Pause / Resume#

POST /api/runs/{id}/pause
POST /api/runs/{id}/resume
Pauses a running run ({"reason": "..."} optional; other body fields are ignored) and resumes it. On resume, a from_rule field is rejected with 400 UNSUPPORTED_FIELD — a paused run continues in place, so re-running from a specific rule has no meaning (see Smart Retry above: cancel the run and retry it instead). A terminal run in either endpoint returns 409 RUN_NOT_ACTIVE.

Logs#

GET /api/runs/{id}/logs
Returns the execution log text — bounded to the newest 8 MiB with an explicit [log truncated …] marker line when the log is larger (issue #734; buffering a multi-GB log per request was a memory-amplification surface). Read the full file from the run workdir when the marker appears.

Results#

GET /api/runs/{id}/results
Returns output file tree with sizes and types.

Files — download / preview / zip#

GET /api/runs/{id}/files?path=<relative-path>
GET /api/runs/{id}/files?path=<relative-path>&preview=true
The read-only result-delivery layer:

  • file → bytes with ETag, Content-Disposition: attachment, and single-range support (Range: bytes=a-b → 206; malformed/multi-range requests degrade to the full body per RFC 9110)
  • directory → a streaming STORE-mode zip (no temporary archive)
  • preview=true → truncated JSON for text-ish formats (100 KB cap) or inline image bytes; other types return 415 NO_PREVIEW
  • paths are sandboxed to the run's workdir (traversal rejected); sensitive filenames (.env, keys, credentials) are never served

Instances#

GET /api/runs/{id}/instances
The sample×rule instance table: every expanded instance the checkpoint knows about (qc_auto-discovered_S1 → rule qc, group auto-discovered, sample S1) with status, duration, and exit code — answers "which sample under which rule failed".

Upload & list user inputs#

POST /api/files        # multipart: field "path" (optional subdir) + file parts
GET  /api/files        # list the acting user's uploaded inputs
Uploads land in workspace/users/<user>/inputs/ (chunked to disk, 8 GiB per-file cap).


Data Discovery#

Analyze Data#

POST /api/data/analyze
Content-Type: application/json

{"paths": ["/data/*.fastq.gz", "/data/*.bam"], "max_depth": 2}
Deterministic file scanning + format inference + pipeline recommendation. Returns { files: [{ path, size, format, format_confidence, paired_with?, sample_name? }], summary, suggested_workflow }. Format detection uses filename extension matching — not AI.

Reference Discovery#

POST /api/data/reference
Content-Type: application/json

{"genome": "hg38", "components": ["fasta", "gtf", "star_index"]}
Finds installed reference genome components and reports missing ones with download commands.

Reference Status#

GET /api/data/reference/status
Which common reference genomes (hg38, mm10, …) have their expected components (fasta, STAR index, GTF) installed at the standard /data/references/ locations. Returns { genome: { components: { component: present } } } per entry — the deterministic counterpart to the POST discovery above.

Data Perception#

POST /api/data/perceive
Content-Type: application/json

{"paths": ["/data/"], "description": "RNA-seq paired-end human samples"}
Data-profiling report built from paths (the same scanner as /api/data/analyze) or from a natural-language description. Powers the "what data do I have" panel.

Samplesheet Parsing#

POST /api/data/samplesheet/parse
Content-Type: application/json

{"content": "sample,fastq_r1,fastq_r2,condition\nS1,a_R1.fastq.gz,a_R2.fastq.gz,ctrl"}
Parses pasted CSV/TSV samplesheet content into the standard column structure (sample, fastq_r1, fastq_r2, condition). Returns { format, samples: [...] } — the bridge from a spreadsheet to [[sample_groups]].


Templates#

GET    /api/templates
POST   /api/templates
GET    /api/templates/{id}
DELETE /api/templates/{id}
Built-in and user-created pipeline templates. System templates are read-only. The list endpoint does not accept filter query parameters.


Plugins#

Validate Plugin#

POST /api/plugins/validate
Content-Type: application/json

{"manifest": {"name": "...", "version": "1.0", "plugin_type": "rule"}, "trusted_keys": {"key1": "hex..."}}
Validates a plugin manifest and optionally verifies its HMAC signature against trusted keys. Returns { valid, name, version, plugin_type, signature_valid, errors }.


AI (Phase 2 — calls deterministic APIs above)#

POST /api/ai/translate          # NL intent → validated .oxoflow (JSON response)
POST /api/ai/translate/stream   # Same, streamed over SSE (progress → done events)
POST /api/ai/explain            # Explain run failure + suggest fix (run owner or admin)
POST /api/ai/interpret          # Interpret results with caveats (run owner or admin)
POST /api/ai/optimize           # Optimize pipeline parameters (pipeline reader via id; owner-supplied TOML unaffected)
GET  /api/ai/config             # Get AI provider configuration (public)
POST /api/ai/config             # Update the shared provider (admin-only outside personal mode)
POST /api/ai/test               # Test the provider (admin-only outside personal mode)
GET  /api/ai/config/user        # The acting user's private provider config
PUT  /api/ai/config/user        # Set the acting user's private provider config
GET  /api/ai/config/server      # Server-level provider config (admin-only outside personal mode)
PUT  /api/ai/config/server      # Set the server-level provider config
GET  /api/ai/config/effective   # The config a run would use: user → server → env fallback chain
GET  /api/knowledge/tools       # Tool catalog the AI can call
GET  /api/knowledge/skills      # Skill catalog the AI can follow

Both knowledge endpoints accept ?q=<query>&limit=<1-50> (default 20). /api/knowledge/tools treats an empty q as a browse: it returns the first limit Bioconda entries. /api/knowledge/skills tokenizes the query and ranks matches (rare terms first); an empty or unmatched q returns an empty skills array — total is always the full catalog size, not the hit count.

The translate endpoints take { "intent": "<natural-language description>", "context": { "data_analysis_id": "<uuid>" } } — intent is REQUIRED (sending only a prompt field fails deserialization); context is optional.

The config endpoints form a three-level resolution chain: a user's private provider overrides the server-wide one, which overrides environment variables. GET /api/ai/config/effective shows the resolved result and which level supplied each field. API keys are stored encrypted at rest (#205) and never echoed back — responses mark key_set instead.

See AI Translation Layer for details.


Chat (v0.8 AI Companion)#

POST /api/chat/send        # Streamed chat over SSE
POST /api/chat/send/json   # Same, single JSON response
GET  /api/chat/sessions    # The acting user's chat sessions

OAuth#

POST /api/auth/oauth/authorize   # Begin the OAuth flow → provider authorize URL
POST /api/auth/oauth/callback    # Provider callback → session token

The redirect URI is OXO_FLOW_OAUTH_REDIRECT_URI when set, with a fixed fallback of http://localhost:3000/api/auth/oauth/callback (see the environment table in AGENTS.md).


Collaboration (Phase 3)#

POST /api/pipelines/{id}/fork    # Fork into workspace (owner = session user)
POST /api/pipelines/{id}/share   # Share (link or workspace; URL uses the bound port)
POST /api/pipelines/import       # Import from oxo+https:// URL
GET  /api/share/{token}          # PUBLIC landing payload (no session required)
GET /api/share/{token} powers the share landing page: pipeline identity, DAG rule order, TOML, owner, expiry, and the most recent terminal run — the token itself is the authorization. Expired links return 410.

See Collaboration for details.

Version History#

GET  /api/pipelines/{id}/revisions        # Snapshot list (newest first, ≤ 50)
GET  /api/pipelines/{id}/revisions/{rev}  # One snapshot's full TOML
POST /api/pipelines/{id}/rollback         # {"revision_id": ...} — restore
Every save/update snapshots the previous content; rollback preserves the current version as a new revision (nothing is lost).

Run Administration#

POST /api/runs/{id}/clean              # CLI clean on the run's workdir
POST /api/runs/{id}/resume-checkpoint  # CLI resume from .oxo-flow/checkpoint.json
resume-checkpoint continues an unfinished run in place as a NEW run row ({"max_jobs": 2} optional). Both are ownership-checked like every other run endpoint.

Webhooks#

GET /api/webhook   # admin-only outside personal mode (#519) — { enabled, url, secret_set, events, signature_scheme }, secret never echoed
PUT /api/webhook   # admin-only outside personal mode — {"enabled": true, "url": "https://...", "secret": "...", "events": [...], "signature_scheme": "..."}
Runs POST a signed payload to the configured URL on terminal states. The signature scheme is configurable:

  • sha256-keyed (default) — sha256(secret‖body), the pre-v0.12 format. The default keeps existing webhook consumers working across upgrades.
  • hmac-sha256 — RFC 2104 HMAC (X-OxoFlow-Signature: hmac-sha256=<hex>), the recommended opt-in for new deployments.

Admin-only outside personal mode — the endpoint is shared infrastructure.

API Keys#

POST   /api/auth/keys       # {"name": "ci-bot"} → { id, name, key } (shown once)
GET    /api/auth/keys       # The acting user's keys (hashes only)
DELETE /api/auth/keys/{id}  # Revoke immediately
Machine credentials: send X-API-Key: oxo_... instead of a Bearer session. Keys resolve to the same ownership context (a key's requests see exactly what its owner sees), are stored as SHA-256 hashes, and revocation is immediate.

Quota#

Runs pre-flight the quota tracker with the workflow's declared threads and memory; over-limit requests get 429 QUOTA_EXCEEDED with the violation list.

GET  /api/quota           # current limits and usage
PUT  /api/quota           # update limits (admin-only; personal mode bypasses admin checks)

PUT /api/quota accepts { max_concurrent_runs, max_total_threads, max_total_memory_mb, max_runs_per_day } — all four fields are required (no defaults applied server-side). Usage is visible at GET /api/quota ({ enabled, limits: {…}, usage: { active_runs, used_threads, used_memory_mb, runs_today } }).

Cluster Connections & Remote Execution#

GET    /api/clusters                  # configured SSH connections
POST   /api/clusters                  # upsert (admin-only outside personal mode)
DELETE /api/clusters/{id}
POST   /api/clusters/{id}/probe       # SSH connectivity + scheduler detection (admin-only outside personal mode)
The stored ssh_key is a secret (#517): it is never returned by any response (including the upsert echo) — clients see ssh_key_set: true | false instead. When OXO_FLOW_MASTER_KEY is set, the key is sealed at rest with the same AES-256-GCM scheme as AI provider keys. POST /api/runs accepts cluster_id: the run then stages its workdir to the remote host (tar over stdio — no rsync), executes under a per-run nohup wrapper, and pulls the results back on completion so every downstream endpoint (logs, files, report) works unchanged. See Run on a cluster.


HPC#

GET /api/hpc
Returns scheduler status (SLURM, PBS/Torque, LSF, SGE), available queues, and node count. This route is only mounted when the server runs in hpc mode.


DAG Editing#

POST /api/pipeline/{id}/command   # Apply an edit command (add/remove rule, connect, ...)
POST /api/pipeline/{id}/undo      # Revert the last edit (JSON body: {"toml_content": ...})
POST /api/pipeline/{id}/redo      # Re-apply the last undone edit (JSON body: {"toml_content": ...})

undo/redo require a JSON body carrying the caller's current TOML; the stacked edit is applied only when it matches that state (otherwise 404). See DAG Edit API for the full command reference.


See Also#