AI Flow API
Reference for the aiflow.svc (AI Flow) HTTP surface. For concepts and authoring see the AI Flow docs.
The HTTP surface of AI Flow (aiflow.svc). Routes are registered in rest/root.go.
- Base URL: your deployment host (e.g.
https://ai-ml.example.com). - Content type:
application/json. - Auth:
Authorization: Bearer <JWT>on most routes; some accept an API key;/admin/*require superadmin. - Authorization: most workflow routes are additionally checked against per-tenant access rules
by permission name (e.g.
workflow.start). A denied call returns 403.
Rate limits
Section titled “Rate limits”Inbound requests are limited per (class, tenant, caller), on by default. The caller identity is part of the bucket, so one runaway client cannot exhaust its tenant’s allowance.
| Class | Routes | Burst | Sustained per second |
|---|---|---|---|
| Read | the GET list, status and config routes | 240 | 60 |
| Write | conversation, request and response turns | 60 | 20 |
| Execute | /workflow/start, /resume, /signal, /cancel | 20 | 5 |
| Admin | /admin/* and /superadmin/* | 30 | 5 |
Execute is the tightest because each of those calls admits a whole workflow run, which spends the deployment’s compute rather than just its request budget.
Buckets are token buckets: burst is what a caller may spend at once, sustained is the refill rate.
Burst sits above sustained deliberately, so a page that fires six reads on load is normal traffic
rather than a 429.
Over the ceiling the route answers 429. Back off and retry; a client that retries immediately
spends its refill as fast as it arrives.
A deployment overrides any class under a ratelimit: boot block. Limits are per process by
default; a multi-replica deployment that wants one shared ceiling sets a distributed store in that
block.
Four families of endpoints share this service:
- Runs: the kind-aware surface. Start, watch, signal and cancel a run of any engine without naming which engine.
- Workflow control: the older DAG-shaped handlers — start/resume/signal/status.
- Operator plane: agents, queue, cross-tenant listings, the data-lifecycle plane.
- Inbound triggers: webhooks, and the record of what they sent.
1. Runs
Section titled “1. Runs”A run is the noun. The engine that owns it — DAG, graph, state machine — is resolved from
the definition, not spelled by the caller. A definition already carries its shape (start: is a
graph, initial: is a state machine), so restating it in a URL would repeat what the file says —
and a caller holding a run id from a log line does not have the kind to restate.
| Method | Path | Purpose |
|---|---|---|
| POST | /runs | Start a run. |
| GET | /runs | List this tenant’s runs, across every engine. |
| GET | /runs/:id | One run’s status, from its own engine. |
| GET | /runs/:id/events | Server-sent events: watch a run to completion. |
| POST | /runs/:id/signal | Move a run that is waiting. |
| POST | /runs/:id/cancel | Stop a run. |
| GET | /engines | What this tenant runs, and what each engine can do. |
POST /runs: start
Section titled “POST /runs: start”Body
| Field | Type | Notes |
|---|---|---|
definitionid | string | Required. |
input | object | Initial input / context. |
engine | string | Optional override. An escape hatch for a definition the service cannot classify, not the normal path. |
Query: wait — optional duration, blocks up to that long for a terminal state, capped at 30s.
curl -X POST https://<host>/runs \ -H "Authorization: Bearer $JWT" -H "Content-Type: application/json" \ -H "X-Customer: acme" -H "X-Env: prod" -H "X-Product: ai-flow" -H "X-Tenant: main" \ -d '{"definitionid": "approval-machine"}'{ "runid": "approval-machine-1790108494503196000", "kind": "statemachine", "live": true, "state": "awaiting:submit", "engine": "statemachine" }state is uniform across engines. A parked run says what would move it —
awaiting:submit — rather than merely “suspended”, because “why is nothing happening” and
“what do I send” are the same question.
GET /runs/:id/events: stream
Section titled “GET /runs/:id/events: stream”text/event-stream. The replacement for polling status in a loop.
event: statedata: {"runid":"…","state":"awaiting:submit","live":true,…}
event: enddata: {"runid":"…","state":"completed","live":false,…}One message per change — a machine resting for an hour costs one frame and a heartbeat, not
twelve thousand identical ones. Heartbeat every 20s as an SSE comment, so a client’s onmessage
never sees it; maximum stream life 30 minutes; ends with event: end at a terminal state.
wait vs events
Section titled “wait vs events”Both exist because they answer different questions. wait= is for a caller that wants one
answer and will hold a request open for it — capped, because an HTTP request is not a
subscription. events is for a caller that wants every change. Neither is a polling loop,
which is what both replaced.
POST /runs/:id/signal: move a parked run
Section titled “POST /runs/:id/signal: move a parked run”{event, payload}. A run not listening for that name is refused 409, and the refusal says
what it is listening for — “not listening” alone sends somebody to check their spelling when
the run has simply moved on. An unknown run is 404.
2. Workflow & agent control
Section titled “2. Workflow & agent control”These are shared cluster handlers mounted by AI Flow. They start, resume, signal, cancel, and inspect workflow instances.
Starting a workflow
Section titled “Starting a workflow”| Method | Path | Permission | Notes |
|---|---|---|---|
| POST | /workflow/start | workflow.start | Async start. Also /product/:product/workflow/start. |
| POST | /workflow/start/sync | workflow.start | Synchronous; optional ?async=. |
| POST | /workflow/start/sync/response | workflow.start | Start and block for responses. |
Start body (all optional unless noted; provide one of workflow/workflowname/workflowid/workflowpath):
| Field | Meaning |
|---|---|
workflowname | Name of a deployed definition. |
workflowid | Definition ID. |
workflow | Inline definition object (same shape as a YAML flow). |
workflowpath | Path/URI to a definition. |
starttask | Node to start at (defaults to the first). |
tasklist | Restrict to a subset of tasks. |
context | Initial context (your vars overrides). |
model | Global model routing hint. |
selectors | Per-task selector overrides: { "task-name": { "gpu": true } }. |
affinity, affinity-key, pin-agent, restart-on-failure | Affinity controls. |
tag / workertag, workerdomain, workerid, workerids | Worker routing. |
curl -X POST https://<host>/workflow/start/sync/response \ -H "Authorization: Bearer $JWT" -H "Content-Type: application/json" \ -d '{ "workflowname": "deployment-pipeline", "context": { "strategy": "canary", "version": "2.1.0" }, "selectors": { "build-one-service": { "tag": "builder" } } }'Resuming, signaling, canceling
Section titled “Resuming, signaling, canceling”| Method | Path | Permission | Body |
|---|---|---|---|
| POST | /workflow/resume, /workflow/event | workflow.resume | instanceid, event, node, data |
| POST | /workflow/signal | workflow.signal | instanceid, signal_name, data: resumes a Wait node |
| POST | /workflow/cancel | workflow.cancel | instanceid |
curl -X POST https://<host>/workflow/signal \ -H "Authorization: Bearer $JWT" -H "Content-Type: application/json" \ -d '{"instanceid":"01HB..","signal_name":"approved","data":{"approver_name":"alice"}}'Status & inspection
Section titled “Status & inspection”| Method | Path | Permission | Purpose |
|---|---|---|---|
| GET | /workflow/status | workflow.status | Instance status. Query: instanceid. |
| GET | /workflow/responses | workflow.status | Responses for an instance. Query: instanceid. |
| GET | /workflow/logs | workflow.logs | Instance logs. |
| GET | /workflow/config | workflow.config | What definitions/tasks the tenant has (API-key or JWT). |
| GET | /workflow/list | workflow.list | List instances. |
Agents & queue
Section titled “Agents & queue”| Method | Path | Permission | Purpose |
|---|---|---|---|
| GET | /agents | agents.list | List registered agents. |
| GET | /workflow/queue | queue.status | Overflow work-queue state (tenant-scoped). |
| DELETE | /workflow/queue/:id | queue.cancel | Cancel a queued item (ownership-checked). |
Superadmin (/admin/*)
Section titled “Superadmin (/admin/*)”Require CheckSuperAdminAccess (cross-tenant visibility):
| Method | Path |
|---|---|
| GET | /admin/agents, /admin/agents/:id |
| GET | /admin/queue, /admin/instances |
| DELETE | /admin/queue/:id |
2a. Inbound triggers
Section titled “2a. Inbound triggers”A trigger is the one way into this service that carries no identity. Every other route arrives with a token and four tenant headers; a webhook arrives from GitHub or Stripe, which have neither and never will.
Declared in configuration, not in code:
triggers: - name: on-push type: webhook # webhook | cron | gitpoll | filesystem active: true config: path: /hooks/github secret: "${GITHUB_WEBHOOK_SECRET}" metadata: engine: dag # which engine kind acts start: deploy # ... and what it runs tenant: acme:prod:ai-flow:main
- name: on-approval type: webhook active: true config: {path: /hooks/approve, secret: "…"} metadata: engine: statemachine signal: submit # move a run that is waiting runfrom: body.run_id # the event field naming WHICH run tenant: acme:prod:ai-flow:mainmetadata keys are a closed set — a typo like singal: is refused at boot rather than
leaving a trigger that fires forever and does nothing.
runfrom: reads the event, not the request body, and for a webhook those differ: the event
is an envelope {body, headers, method, path, query}, so a field posted as {"run_id": …} is at
body.run_id. Dotted paths resolve; a bare run_id finds nothing.
Three properties of a webhook path
Section titled “Three properties of a webhook path”- It is exempt from the tenant gate. Every other request must name a tenant in headers, and an outside sender cannot. The tenant comes from the binding instead.
- Letting the caller choose the tenant would be wrong — an unauthenticated request naming the tenant it acts on. The binding names it in your configuration.
- It is not exempt from authentication. A webhook verifies an HMAC-SHA256 over the raw body
when
secret:is set. One declared without a secret is mounted and logged atWARNon every boot: refusing it would make local development impossible, and mounting it silently would hide an open endpoint.
It is also not rate-limited, for the same reason it is not token-gated: the limiter keys on a
caller a webhook does not have. Each firing costs a database row, so bound it with
max_body_size, require secret: so the writes are attributable, and rate-limit the declared
paths at your gateway.
The firing log
Section titled “The firing log”| Method | Path | Purpose |
|---|---|---|
| GET | /triggers | What this service listens for, and each binding’s status. |
| GET | /triggers/firings | What has actually arrived. Query: limit, trigger, outcome. |
| GET | /triggers/firings/:id | One firing, with its payload. |
Both reads are token-gated, unlike the webhooks themselves: what a service listens for is a map of the service, and a firing payload is the most sensitive thing it stores — whatever an outside system sent, which may include a credential.
| Outcome | Meaning | Answered to the sender |
|---|---|---|
started | A run began. | 200 |
signalled | A waiting run moved. | 200 |
refused | Well formed; nothing to act on — unknown run, run not listening. | 422 |
failed | The engine was asked and errored. | 500 |
The 422/500 split matters: senders retry a 5xx with backoff, so answering 500 to “that run does not exist” turns one bad request into a permanent retry loop against an answer that will never change. A 4xx also carries its reason back, because it describes the caller’s own request.
A refused firing is the one with no other trace. No run starts and no directory row appears — and those are exactly the firings somebody asks about later: “we sent it on Tuesday, did you get it?”
The payload comes back on the single read, not the listing: a page of fifty webhook bodies is
a lot of sensitive data moved for nobody’s benefit. The listing does carry error, because “it
fired and was refused” is useless without why.
Tracing a run back to its cause
Section titled “Tracing a run back to its cause”Two fields on a run’s directory entry close the loop: which trigger by name, and which firing — and therefore what arrived. The second is not redundant: a trigger that fires hourly leaves a dozen candidate causes a day, and matching by timestamp is a guess.
3. Generic entity API (v2 surfaces)
Section titled “3. Generic entity API (v2 surfaces)”AI Flow auto-mounts the platform data layer-api’s v2 REST and Schema surfaces, exposing generic
per-entity CRUD and schema introspection for the aiflow_* entities, backed by the per-tenant
engine. Consult the the platform data layer-api surface docs for the exact routes; these are primarily for tooling
and admin use rather than the flow-authoring path.
4. Response envelopes & errors
Section titled “4. Response envelopes & errors”- Success (native handlers): the v1 envelope
{"data":{"<entity>": row | [rows]}}, preserved for backward compatibility across the v1→v2 migration. - Errors use codes prefixed
ai-flow-…, e.g.:ai-flow-json-request-parse-failedai-flow-request-required-field-not-found(with{{field}})ai-flow-query-parsing-failedai-flow-ctx-tenant-not-found
- Unmatched routes return
404(logged).
5. Request notes
Section titled “5. Request notes”- Tenant required. Every request resolves a CEPT tenant from its JWT/headers; a missing tenant
yields
ai-flow-ctx-tenant-not-found. - Resolving a run to its engine goes through the run directory. A service running more than one engine needs it; without it, “which engine owns this run” has no answer.
See the authoring flows doc for how to write the flows these endpoints run, and concepts for the objects (run, engine, trigger, firing) these endpoints act on.
Authentication
Section titled “Authentication”Every route requires a bearer token:
Authorization: Bearer <token>Authorization
Section titled “Authorization”The non-open /workflow/* routes evaluate a per-tenant access rule, in the same way as
Workflows. Configure an authorization block for every
tenant as part of onboarding.
Errors
Section titled “Errors”| Status | Meaning |
|---|---|
400 | Request parse or validation failure |
401 | Missing or invalid token, as plain text |
403 | An access rule denied the request, as {"message": "access denied"} |
500 | Unexpected server error |