Skip to content
Talk to our solutions team

AI Flow Operations

Installing, configuring, running, and tuning AI Flow (aiflow.svc).

Terminal window
kvm install ai/aiflow.svc

Check the version:

Terminal window
aiflow.svc version

AI Flow is a cluster node built on lib-cluster/boot. A single binary can run as an orchestrator and/or an agent; both gRPC server credential sets are enabled by default.

Terminal window
./dist/aiflow.svc --config /etc/ai-flow/.ai-flow.yaml
  • The service gates its HTTP surface on the first tenant workflow load (the definition-load gate), so a fresh process won’t start runs until at least one tenant’s definitions are loaded.
  • Agent nodes declare which tenants they serve:
Terminal window
./dist/aiflow.svc --tenants "acme:prod:flow:default,globex:prod:flow:default" --exclusive false

Config is loaded via the chassis (bootconfig.Load("ai-flow"), default file .ai-flow.yaml). There is no config file in the repo: supply one at deploy time (or inject a config map for tests). Keys consumed:

KeyPurpose
httpaddressHTTP listen address.
grpcaddressgRPC listen address (orchestrator / agent server).
infrahostInfra/monitoring address advertised by the agent.
capabilitiesML model proxies this node exposes (routing by model).
tagsAgent routing tags (matched by selectors.tag).
KeyPurpose
vault.url, vaultVault connection. DB passwords and secrets resolve from here (literal fallback only in dev).
KeyPurpose
mlproxiesModel-server proxy definitions.
forceproxyrestartRestart proxies on load (default true).
KeyPurpose
orchestrator.queue.maxdepthOverflow work-queue depth. 0 or negative = unlimited.
orchestrator.store.workitemOverflow-queue backend: database (default) or memory.
orchestrator.store.affinityAffinity-binding backend: database (default) or memory.

Persistence backends. In an AI Flow deployment, execution instances and runs are persisted to PostgreSQL (the platform data layer), not in-memory. Use database stores in production so the overflow queue and affinity bindings survive restarts.

These live in code (the agent fabric); changing them requires a build:

ConstantValue
Message chunk size3.5 MiB
Circuit-breaker max failures / reset5 / 30 s
Flow-control max in-flight1000
Per-agent outbound queue1000
Write/requeue timeout5 s
Heartbeat jitter±10%
gRPC keepalive (client)Time=30s, Timeout=10s, PermitWithoutStream
gRPC keepalive (server)+ MinTime=5s
  • Each tenant maps to its own PostgreSQL schema (resolved via chassis config). Connections set app.tenant_id (and userid) session variables so row-level security applies.
  • The entity model (aiflow_*) is compiled into the binary from embedded entity YAML and built into a process-wide schema at boot. Malformed entity YAML is fatal at startup: catch it in CI.
  • Credential rotation is handled live: TenantConfigChanged drops the affected tenant’s cached engines/pools; the schema is untouched.

Workflow definitions are not deployed with the binary. They’re read per tenant from the tenant’s product: every YAML file under its ai/flows/ folder, through the tenant’s meta client. When a file under ai/flows/ changes, AI Flow re-parses and rebuilds that tenant’s definitions and engine.

To ship a new flow: add or update its YAML under ai/flows/ in the product repository; the change reaches the service by itself. Verify with GET /workflow/config (lists the tenant’s definitions/tasks).

A box with no meta server has no definition source at all, which means no definition provider and no execution engine for any tenant. Every workflow route answers “not found”. For that case the service config can name a directory of workflow YAMLs:

workflows:
localdir: workflows # relative paths resolve against the bootstrap file
loadtenants:
- default:dev:ai-flow:default

Every .yaml/.yml directly under localdir is parsed at boot and published to each tenant key in loadtenants, full four-part customer:env:product:tenant keys, because that is what the engine and every request-path lookup key by. Definitions go through the same refresh path as the product’s ai/flows/, so behaviour after load is identical.

Leave localdir unset in any deployment: it is a development stand-in, not a second supported delivery channel, and it does not re-read the directory after boot.

A worked example, dev-sample.yaml, uses only control-flow nodes (assign, succeed), which the engine runs in-process, a definition with ordinary task nodes still needs a connected agent to make progress.

  • Readiness is effectively “first workflow load complete”: the HTTP surface comes up only once that gate opens.
  • Use GET /workflow/config (per tenant) as a functional check that definitions are loaded, and GET /agents to confirm agents are connected.
  • GET /workflow/queue shows overflow depth; a persistently non-empty queue means insufficient agent capacity for the offered load.

On by default, per (class, tenant, caller). Four classes, sized by what a call costs: read 240/60, write 60/20, execute 20/5 and admin 30/5, as burst over sustained-per-second. Over the ceiling the route answers 429. See the API page for the full table.

Override any class under a ratelimit: boot block. Limits are held per process by default; a multi-replica deployment that wants one shared ceiling configures a distributed store there.

A limiter that fails to build leaves the service serving without inbound limits and says so in the boot log. Check for that line after a config change: an orchestrator that cannot rate-limit should still serve, but you want to know.

The authoritative operator doc is the agent fabric’s tuning guide. Key knobs and signals:

SymptomLikely causeLever
Growing overflow queueNot enough agents / wrong tagsAdd agents; check selectors match agent tags/models.
selector.fail.count rising with NoMatchingAgentFlows request a tag/model no agent hasFix the selector or deploy an agent with that capability.
Agents flapping unhealthy under loadHeartbeats starvedAlready mitigated by streamSendMu; verify heartbeat interval vs. task duration.
backpressure.count climbingFlow-control cap hitMore agents, or reduce fan-out max_concurrency.
Slow selection (selector.duration)Large agent fleet / churnInvestigate heartbeat volume; consider agent sharding by tenant.
  • Heartbeat interval: shorter = faster failure detection, more traffic. Staleness threshold is 2× the interval.
  • max_concurrent_tasks per agent (0 = unlimited) bounds per-agent load and drives the overflow queue.
  • Fan-out: cap Map max_concurrency so a single flow can’t monopolize the fleet.
  • Orchestrator (kis.agent): connected, connect.count, disconnect.count, heartbeat.count, state_change.count, backpressure.count.
  • Selector (kis.selector): duration, success.count, fail.count.
  • Agent (OTel gauges): CPU/memory/GPU/disk usage, data.inflow/outflow.
  • the flow engine EventLog: append-only audit (AffinityBound, AgentFailover, SignalReceived, DataPipeRecord, …) for replay.
  • Instance logs: GET /workflow/logs, and engine GetInstanceLog / GetInstanceRunsLog / GetInstanceMLResponseLog.

All metric registries fall back to no-op instruments, so metric-registration failure never panics the hot path.

Every engine operation is timed. An operation at or above 250 ms earns a log line, and so does any failure, so a read that has gone slow shows up without anything being switched on. A healthy read logs nothing.

The statement log is the second half, and it is off by default because generated SQL names columns. Turn it on from the boot block:

sqllog:
mode: slow # off (default) | slow | all
slower_than: 250ms
sample: 1
max_length: 4096

It also changes at runtime, without a restart, so it can be turned on for a minute during an incident:

Terminal window
curl -X PUT -H "Authorization: Bearer $ADMIN" -H "Content-Type: application/json" \
-H 'X-Customer: acme' -H 'X-Product: shop' -H 'X-Env: prod' -H 'X-Tenant: main' \
-d '{"mode":"all","sample":0.1}' https://<host>/admin/data/sqllog

GET on the same route reads the setting in force along with seen and logged counters. See Data observability for what a line contains and why bound values are never in it.

Entities that declare an audit: block write to the tenant’s own durable, hash-chained data_audit_events table, alongside the access denials the compiler records. One stream, one order, one chain per tenant.

A mutation’s audit row is written on that mutation’s own transaction, as its last statement, so a write that rolls back leaves no row claiming it happened. That matters on this service in particular: submitting a request turn rolls back when the dispatch is refused.

The chain head lives in the table, so a restart or a second replica continues the chain rather than starting a second one. Read it with SQL and verify it by walking the rows in seq order. See Data observability for the verification procedure.

This is separate from the flow engine’s own EventLog above, which records orchestration events for replay rather than access to data.

  • OnShutdown calls the v2boot Shutdown, invalidating all cached engines and closing pools.
  • The orchestrator Drain(timeout) marks agents draining and broadcasts a drain control message; the flow engine performs a 3-phase drain that completes in-flight work before exit.

The service’s entities run on the platform’s data engine and carry its baseline chain, so the schema’s declarations take effect on every read and write:

DeclarationEffect
transforms: on a fieldApplied on write
A compliance directive such as piiApplied on read, so a tagged field is masked for callers without the reveal
references: between entitiesEnforced, so a turn cannot name a conversation that does not exist
audit: on an entityWritten to the tenant’s durable audit stream

See Field compliance directives and Hooks and the execution pipeline.

  • gRPC mTLS configured (both server and client certs present): otherwise it falls back to insecure with a warning. Confirm certs load (cacert/servercert/serverkey [+ clientcert/clientkey]).
  • Vault reachable; no literal DB passwords in production config.
  • Per-tenant access rules configured for the granular permissions (workflow.start, workflow.resume, workflow.signal, workflow.cancel, workflow.status, workflow.logs, workflow.config, workflow.list, agents.list, queue.status, queue.cancel).
  • Superadmin routes (/admin/*) restricted appropriately.
  • RLS verified: cross-tenant reads return nothing.

GET /workflow/list takes its selector as filter= or id=.

See the design spec for the full architecture and orchestration for the agent fabric internals.