AI Flow Operations
Installing, configuring, running, and tuning AI Flow (aiflow.svc).
Install
Section titled “Install”kvm install ai/aiflow.svcCheck the version:
aiflow.svc versionAI Flow is a cluster node built on lib-cluster/boot. A single binary can run as an
orchestrator and/or an agent; both gRPC server credential sets are enabled by default.
./dist/aiflow.svc --config /etc/ai-flow/.ai-flow.yaml- The service gates its HTTP surface on the first tenant workflow load (the definition-load gate), so a fresh process won’t start runs until at least one tenant’s definitions are loaded.
- Agent nodes declare which tenants they serve:
./dist/aiflow.svc --tenants "acme:prod:flow:default,globex:prod:flow:default" --exclusive falseConfiguration
Section titled “Configuration”Config is loaded via the chassis (bootconfig.Load("ai-flow"), default file .ai-flow.yaml).
There is no config file in the repo: supply one at deploy time (or inject a config map for
tests). Keys consumed:
Service / infra
Section titled “Service / infra”| Key | Purpose |
|---|---|
httpaddress | HTTP listen address. |
grpcaddress | gRPC listen address (orchestrator / agent server). |
infrahost | Infra/monitoring address advertised by the agent. |
capabilities | ML model proxies this node exposes (routing by model). |
tags | Agent routing tags (matched by selectors.tag). |
Vault & secrets
Section titled “Vault & secrets”| Key | Purpose |
|---|---|
vault.url, vault | Vault connection. DB passwords and secrets resolve from here (literal fallback only in dev). |
Model proxies
Section titled “Model proxies”| Key | Purpose |
|---|---|
mlproxies | Model-server proxy definitions. |
forceproxyrestart | Restart proxies on load (default true). |
Orchestrator / queue
Section titled “Orchestrator / queue”| Key | Purpose |
|---|---|
orchestrator.queue.maxdepth | Overflow work-queue depth. 0 or negative = unlimited. |
orchestrator.store.workitem | Overflow-queue backend: database (default) or memory. |
orchestrator.store.affinity | Affinity-binding backend: database (default) or memory. |
Persistence backends. In an AI Flow deployment, execution instances and runs are persisted to PostgreSQL (the platform data layer), not in-memory. Use
databasestores in production so the overflow queue and affinity bindings survive restarts.
Fixed constants
Section titled “Fixed constants”These live in code (the agent fabric); changing them requires a build:
| Constant | Value |
|---|---|
| Message chunk size | 3.5 MiB |
| Circuit-breaker max failures / reset | 5 / 30 s |
| Flow-control max in-flight | 1000 |
| Per-agent outbound queue | 1000 |
| Write/requeue timeout | 5 s |
| Heartbeat jitter | ±10% |
| gRPC keepalive (client) | Time=30s, Timeout=10s, PermitWithoutStream |
| gRPC keepalive (server) | + MinTime=5s |
Database & migrations
Section titled “Database & migrations”- Each tenant maps to its own PostgreSQL schema (resolved via chassis config). Connections set
app.tenant_id(anduserid) session variables so row-level security applies. - The entity model (
aiflow_*) is compiled into the binary from embedded entity YAML and built into a process-wide schema at boot. Malformed entity YAML is fatal at startup: catch it in CI. - Credential rotation is handled live:
TenantConfigChangeddrops the affected tenant’s cached engines/pools; the schema is untouched.
Workflow definitions (deploying flows)
Section titled “Workflow definitions (deploying flows)”Workflow definitions are not deployed with the binary. They’re read per tenant from the
tenant’s product: every YAML file under its ai/flows/ folder, through the tenant’s meta client.
When a file under ai/flows/ changes, AI Flow re-parses and rebuilds that tenant’s definitions and
engine.
To ship a new flow: add or update its YAML under ai/flows/ in the product repository; the change
reaches the service by itself. Verify with GET /workflow/config (lists the tenant’s
definitions/tasks).
Local definitions (development only)
Section titled “Local definitions (development only)”A box with no meta server has no definition source at all, which means no definition provider and no execution engine for any tenant. Every workflow route answers “not found”. For that case the service config can name a directory of workflow YAMLs:
workflows: localdir: workflows # relative paths resolve against the bootstrap fileloadtenants: - default:dev:ai-flow:defaultEvery .yaml/.yml directly under localdir is parsed at boot and published to each tenant key
in loadtenants, full four-part customer:env:product:tenant keys, because that is what the
engine and every request-path lookup key by. Definitions go through the same refresh path as
the product’s ai/flows/, so behaviour after load is identical.
Leave localdir unset in any deployment: it is a development stand-in, not a second supported
delivery channel, and it does not re-read the directory after boot.
A worked example, dev-sample.yaml, uses only control-flow nodes (assign, succeed), which the
engine runs in-process, a definition with ordinary task nodes still needs a connected agent to
make progress.
Health & readiness
Section titled “Health & readiness”- Readiness is effectively “first workflow load complete”: the HTTP surface comes up only once that gate opens.
- Use
GET /workflow/config(per tenant) as a functional check that definitions are loaded, andGET /agentsto confirm agents are connected. GET /workflow/queueshows overflow depth; a persistently non-empty queue means insufficient agent capacity for the offered load.
Inbound rate limits
Section titled “Inbound rate limits”On by default, per (class, tenant, caller). Four classes, sized by what a call costs: read 240/60,
write 60/20, execute 20/5 and admin 30/5, as burst over sustained-per-second. Over the ceiling the
route answers 429. See the API page for the full table.
Override any class under a ratelimit: boot block. Limits are held per process by default; a
multi-replica deployment that wants one shared ceiling configures a distributed store there.
A limiter that fails to build leaves the service serving without inbound limits and says so in the boot log. Check for that line after a config change: an orchestrator that cannot rate-limit should still serve, but you want to know.
Tuning
Section titled “Tuning”The authoritative operator doc is the agent fabric’s tuning guide. Key knobs and signals:
| Symptom | Likely cause | Lever |
|---|---|---|
| Growing overflow queue | Not enough agents / wrong tags | Add agents; check selectors match agent tags/models. |
selector.fail.count rising with NoMatchingAgent | Flows request a tag/model no agent has | Fix the selector or deploy an agent with that capability. |
| Agents flapping unhealthy under load | Heartbeats starved | Already mitigated by streamSendMu; verify heartbeat interval vs. task duration. |
backpressure.count climbing | Flow-control cap hit | More agents, or reduce fan-out max_concurrency. |
Slow selection (selector.duration) | Large agent fleet / churn | Investigate heartbeat volume; consider agent sharding by tenant. |
- Heartbeat interval: shorter = faster failure detection, more traffic. Staleness threshold is
2×the interval. max_concurrent_tasksper agent (0= unlimited) bounds per-agent load and drives the overflow queue.- Fan-out: cap Map
max_concurrencyso a single flow can’t monopolize the fleet.
Observability
Section titled “Observability”- Orchestrator (
kis.agent):connected,connect.count,disconnect.count,heartbeat.count,state_change.count,backpressure.count. - Selector (
kis.selector):duration,success.count,fail.count. - Agent (OTel gauges): CPU/memory/GPU/disk usage,
data.inflow/outflow. - the flow engine EventLog: append-only audit (
AffinityBound,AgentFailover,SignalReceived,DataPipeRecord, …) for replay. - Instance logs:
GET /workflow/logs, and engineGetInstanceLog/GetInstanceRunsLog/GetInstanceMLResponseLog.
All metric registries fall back to no-op instruments, so metric-registration failure never panics the hot path.
Slow operations and the statement log
Section titled “Slow operations and the statement log”Every engine operation is timed. An operation at or above 250 ms earns a log line, and so does any failure, so a read that has gone slow shows up without anything being switched on. A healthy read logs nothing.
The statement log is the second half, and it is off by default because generated SQL names columns. Turn it on from the boot block:
sqllog: mode: slow # off (default) | slow | all slower_than: 250ms sample: 1 max_length: 4096It also changes at runtime, without a restart, so it can be turned on for a minute during an incident:
curl -X PUT -H "Authorization: Bearer $ADMIN" -H "Content-Type: application/json" \ -H 'X-Customer: acme' -H 'X-Product: shop' -H 'X-Env: prod' -H 'X-Tenant: main' \ -d '{"mode":"all","sample":0.1}' https://<host>/admin/data/sqllogGET on the same route reads the setting in force along with seen and logged counters. See
Data observability for what a line
contains and why bound values are never in it.
The audit stream
Section titled “The audit stream”Entities that declare an audit: block write to the tenant’s own durable, hash-chained
data_audit_events table, alongside the access denials the compiler records. One stream, one order,
one chain per tenant.
A mutation’s audit row is written on that mutation’s own transaction, as its last statement, so a write that rolls back leaves no row claiming it happened. That matters on this service in particular: submitting a request turn rolls back when the dispatch is refused.
The chain head lives in the table, so a restart or a second replica continues the chain rather than
starting a second one. Read it with SQL and verify it by walking the rows in seq order. See
Data observability for the verification procedure.
This is separate from the flow engine’s own EventLog above, which records orchestration events for replay rather than access to data.
Graceful shutdown
Section titled “Graceful shutdown”OnShutdowncalls the v2bootShutdown, invalidating all cached engines and closing pools.- The orchestrator
Drain(timeout)marks agents draining and broadcasts a drain control message; the flow engine performs a 3-phase drain that completes in-flight work before exit.
What the engine applies to flow data
Section titled “What the engine applies to flow data”The service’s entities run on the platform’s data engine and carry its baseline chain, so the schema’s declarations take effect on every read and write:
| Declaration | Effect |
|---|---|
transforms: on a field | Applied on write |
A compliance directive such as pii | Applied on read, so a tagged field is masked for callers without the reveal |
references: between entities | Enforced, so a turn cannot name a conversation that does not exist |
audit: on an entity | Written to the tenant’s durable audit stream |
See Field compliance directives and Hooks and the execution pipeline.
Security checklist
Section titled “Security checklist”- gRPC mTLS configured (both server and client certs present): otherwise it falls back to
insecure with a warning. Confirm certs load (
cacert/servercert/serverkey[+clientcert/clientkey]). - Vault reachable; no literal DB passwords in production config.
- Per-tenant access rules configured for the granular permissions (
workflow.start,workflow.resume,workflow.signal,workflow.cancel,workflow.status,workflow.logs,workflow.config,workflow.list,agents.list,queue.status,queue.cancel). - Superadmin routes (
/admin/*) restricted appropriately. - RLS verified: cross-tenant reads return nothing.
Filtering workflow requests and responses
Section titled “Filtering workflow requests and responses”GET /workflow/list takes its selector as filter= or id=.
See the design spec for the full architecture and orchestration for the agent fabric internals.