Entity reference
AI Flow persists its state as entities on the platform’s Data engine, in the tenant’s own schema. They are ordinary entities: the same access rules, row-level security, compliance directives and audit stream apply, and they are reachable through the generic entity API as well as through this service’s own routes.
Eleven entities, all prefixed aiflow_.
Each engine keeps its own run table
Section titled “Each engine keeps its own run table”This is the organising idea, and the reason there is no single “runs” table.
A DAG run is a set of nodes with dependencies plus a row per node attempt. A graph run is a position and a step count. A state-machine run is a state name and a context map. Those are genuinely different shapes, and one table holding all three would be a union of columns where most are null for any given row.
So each engine persists what it has, and a directory sits over them:
| Engine | Writes | Holds |
|---|---|---|
| dag | aiflow_run, aiflow_task | A run, and one row per node attempt |
| graph | aiflow_graph_run | Current node, step count, the path walked |
| statemachine | aiflow_machine_run | Current state, context, history |
| datapipe | not here | Pipes belong to datapipe.svc, under its own datapipeline_ prefix |
| all of them | aiflow_engine_run | The directory: which engine owns which run |
aiflow_engine_run is what makes “I have a run id — which engine ran it?” answerable. It holds no
state of its own, so it cannot disagree with an engine about one, and the engine that owns a run
writes its directory row on the same save that writes its own state. One writer, no drift.
Which library declares which table
Section titled “Which library declares which table”The tables are not all AI Flow’s. Two libraries beneath it contribute entity fragments and the service composes them at boot; AI Flow itself declares none.
That matters when you are reading a schema and wondering where a column’s meaning is defined, and it is why every service that registers engines gets the same directory and queue tables.
| Declared by | Entities | Why it lives there |
|---|---|---|
| lib-cluster (the substrate) | aiflow_engine_run, aiflow_trigger_firing, aiflow_engine, aiflow_agent, aiflow_work_item, aiflow_affinity_binding | Every service registering engines needs a run directory, a firing log, a work queue and an agent registry. Declaring them once is why they agree across services. |
| lib-guppi (the engines) | aiflow_run, aiflow_task, aiflow_graph_run, aiflow_machine_run, aiflow_checkpoint | Engine-specific run state — each engine’s own shape. |
| AI Flow | — | Contributes no fragment. Everything it stores belongs to a layer below. |
The aiflow_ prefix is applied by the datastore rather than baked into the fragments, which is how
four services compose the same tables under four different prefixes.
The shape of it
Section titled “The shape of it” aiflow_engine_run ── the directory over every engine's runs │ ├── dag ─────────── aiflow_run ──< aiflow_task │ a flow run one node attempt │ └──< aiflow_checkpoint │ iterator progress ├── graph ───────── aiflow_graph_run └── statemachine ── aiflow_machine_run
aiflow_trigger_firing what a trigger sent, and what it did aiflow_work_item a queued task >── aiflow_agent aiflow_affinity_binding a key pinned >── a registered worker aiflow_engine an engine's livenessThe agent relations join on name, not id. aiflow_agent has both — id is a generated
ULID, name is the logical agent id (harness-agent-1) — and every other table stores the
name: aiflow_work_item.agent_id, aiflow_affinity_binding.agent_id and
aiflow_task.workerid all hold harness-agent-1, never the ULID.
Columns every entity carries
Section titled “Columns every entity carries”All eleven inherit kisai.common, so these six exist on every table without being declared:
| Column | Type | Nullable | Set by |
|---|---|---|---|
createdby | string | no | The caller’s identity, on create. Final |
createdon | timestamp | no | Creation time. Final |
updatedby | string | yes | The caller’s identity, on each write |
updatedon | timestamp | yes | Last write time |
deletedby | string | yes | Soft delete |
deletedon | timestamp | yes | Soft delete |
createdby is load-bearing rather than decorative on the user-facing tables: reads are scoped to
createdby = <caller>, built structurally by the handler rather than written as a rule.
Conventions
Section titled “Conventions”Ordering. Every entity a person lists is createdon desc, id desc, so a list is newest-first.
The exceptions order by what they are dispatched or addressed by: aiflow_agent by name,
aiflow_work_item by priority, createdon, aiflow_affinity_binding by createdon,
aiflow_checkpoint by run_id, node_name, and aiflow_trigger_firing by fired_on desc.
Timestamps are of two kinds, deliberately. createdon and its siblings are real timestamptz
columns from kisai.common. Columns an engine writes itself — starttime, endtime, ended_on,
fired_on, lease_expires_at — are bigint epoch nanoseconds, because they are compared and
arithmetic’d on the hot path.
Reading one as the other is the common mistake, and it fails quietly: a timestamptz read as an
integer yields 0, which renders as year one rather than as an error.
Audit. An entity records nothing until its audit: block says so. Where it is on, field names
and primary keys are recorded, never values. On this service that is aiflow_run, aiflow_task
and aiflow_agent. See Operations.
Access. The engine denies by default, so an operation with no rule is refused for every external caller. Each entity opts in only to what this service’s own handlers perform, and the service writes execution state under its own service context, which is not gated.
Eight entities carry no access block at all, which IS the closed posture: aiflow_engine_run,
aiflow_trigger_firing, aiflow_engine, aiflow_agent, aiflow_work_item,
aiflow_affinity_binding, aiflow_checkpoint, and the two native-engine run tables
aiflow_graph_run and aiflow_machine_run. They are the orchestrator’s own bookkeeping; a caller
reaches what they hold through /runs, /triggers and the agent routes rather than by reading
rows.
The actions block is sealed (access-lock) throughout. A product layer composed on top may
not replace it and grant itself operations this service refused. It may still add rls or fields
rules, which can only narrow what a caller reaches.
The directory
Section titled “The directory”aiflow_engine_run
Section titled “aiflow_engine_run”Which engine owns which run. One row per run, whatever engine ran it, and no state — asked for a run’s detail, the service asks the engine that owns it.
Declared by lib-cluster. Written by the engine on the same save as its own run row; a trigger fills in cause. Recording merges rather than replaces, so whichever fact arrives second completes the row instead of erasing the other’s half.
| Field | Type | Required | Description |
|---|---|---|---|
id | ulid | yes | Primary key. Final |
run_id | string | yes | The engine’s own run id |
engine_kind | string | yes | dag, graph, statemachine, datapipe, test |
tenant_key | string | yes | The full tenant key |
definition_id | string | no | What was run |
parent_run_id | string | no | Set when another run spawned this one |
triggered_by | string | no | Which trigger caused it, by name. Empty when a caller started it |
trigger_firing_id | string | no | Which firing — and therefore what arrived |
ended_on | bigint | no | Epoch ns. Null while live; first write wins |
triggered_by and trigger_firing_id are not redundant. The first says a trigger caused this run;
the second says which firing, so the payload is reachable. A trigger that fires hourly leaves a
dozen candidate causes a day, and matching them by timestamp is a guess.
Indexes. aiflow_engine_run_tenant_run_uq, unique on (tenant, run), which is what makes
recording idempotent — an engine writes the directory on every save, so a create-only write would
fail on the second step of every run. aiflow_engine_run_tenant_created serves the newest-first
listing.
The row stays after a run ends. “Which engine ran this” is asked about finished runs too, by exactly the operator the directory is for.
Per engine
Section titled “Per engine”aiflow_run — the DAG engine
Section titled “aiflow_run — the DAG engine”One execution of a DAG definition. Declared by lib-guppi.
| Field | Type | Required | Description |
|---|---|---|---|
id | ulid | yes | Primary key. Final |
definitionid | string | yes | The definition executed |
runtype | string | yes | dag. Historical rows may hold datapipeline |
context | object | no | The run’s shared data structure |
metadata | object | no | Worker assignments, tasklist, affinity key, the execution plan |
states | object | no | State history; the last entry is the current state |
starttime | bigint | no | Epoch ns |
endtime | bigint | no | Epoch ns |
duration | bigint | no | endtime - starttime |
Indexes. aiflow_run_owner_recent_idx on (createdby, createdon) — the ownership predicate is
unconditional on every user-facing read; aiflow_run_definition_idx on definitionid.
Audited. Access: actions sealed.
aiflow_task — the DAG engine
Section titled “aiflow_task — the DAG engine”One node-execution attempt within a run, bound to a worker. Declared by lib-guppi.
| Field | Type | Required | Description |
|---|---|---|---|
id | ulid | yes | Primary key. Final |
name | string | yes | The node name. A fan-out attempt carries a synthetic suffix |
definitionid | string | yes | Denormalised, for the definition index |
runid | string | yes | Parent run — the run_task relation |
workerid | string | no | The agent name that executed it, not a ULID |
context | object | no | The task’s own data |
metadata | object | no | Parallel group, pending events |
states | object | no | State history |
starttime / endtime / duration | bigint | no | Epoch ns |
resumedata | object | no | What a suspended node needs in order to continue |
Indexes. aiflow_task_run_node_idx on (runid, name); aiflow_task_definition_idx.
Audited. Access: actions sealed.
aiflow_graph_run — the graph engine
Section titled “aiflow_graph_run — the graph engine”A graph may cycle, so a run is a position plus a bound. Declared by lib-guppi.
| Field | Type | Required | Description |
|---|---|---|---|
id | ulid | yes | Primary key. Final |
run_id | string | yes | Unique. The engine’s own id |
definition_id | string | yes | What is being walked |
task | string | no | The current node |
steps | bigint | no | Steps taken, checked against the definition’s maxsteps |
path | text | no | The nodes walked, in order |
run_counts | text | no | Per-node visit counts — a cycle’s bookkeeping |
state | text | no | The run’s own data |
status | string | no | Engine status |
reason | string | no | Why it stopped |
error | text | no | Failure detail |
Index. aiflow_graph_run_runid_uq, unique on run_id.
Not audited. No access block — closed.
aiflow_machine_run — the state-machine engine
Section titled “aiflow_machine_run — the state-machine engine”A machine rests between events, so a run is a state and what it carries. Declared by lib-guppi.
| Field | Type | Required | Description |
|---|---|---|---|
id | ulid | yes | Primary key. Final |
run_id | string | yes | Unique |
definition_id | string | yes | The machine |
state | string | no | The state it is resting in |
context | text | no | The machine’s data |
history | text | no | States visited, in order |
live | bool | no | False once terminal |
suspended | bool | no | Waiting on something other than an event |
failure | text | no | Failure detail |
reason | string | no | Why it stopped |
Index. aiflow_machine_run_runid_uq, unique on run_id.
Not audited. No access block — closed.
aiflow_checkpoint
Section titled “aiflow_checkpoint”Iterator progress for a node that fans out, so a resumed run does not repeat work. Declared by
lib-guppi; run_checkpoint is the relation to aiflow_run.
| Field | Type | Required | Description |
|---|---|---|---|
id | ulid | yes | Primary key. Final |
run_id | string | yes | The run |
node_name | string | yes | The fanning node |
definition_id | string | no | Denormalised |
iterator_type | string | no | map, foreach |
cursor | text | no | Where to resume |
idx / total | bigint | no | Position and size |
done | bool | no | Iteration complete |
updated_at | bigint | no | Epoch ns |
Index. aiflow_checkpoint_run_node_uq, unique on (run_id, node_name) — one checkpoint per
fanning node, so a second write updates rather than duplicating.
Not audited. No access block — closed.
Triggers
Section titled “Triggers”aiflow_trigger_firing
Section titled “aiflow_trigger_firing”What an inbound trigger sent, and what became of it. Declared by lib-cluster.
Append-only. Unlike the directory, which an engine rewrites as a run advances, a firing is a fact about one instant: written once, after the outcome is known. There is no update path and no delete path, because an audit trail that can be edited is not one.
| Field | Type | Required | Description |
|---|---|---|---|
id | ulid | yes | Primary key. Final |
tenant_key | string | yes | Every read is qualified by it — an id alone is not a key |
trigger_name | string | yes | Which trigger fired |
trigger_type | string | no | webhook, cron, gitpoll, filesystem |
payload | text | no | What arrived, as JSON. The whole reason the table exists |
action | string | no | start or signal — what the binding meant to do |
definition_id | string | no | For a start |
event | string | no | For a signal |
run_id | string | no | What it started or moved. Empty on a refusal |
outcome | string | no | started, signalled, refused, failed |
error | text | no | Why, on a refusal or a failure |
fired_on | bigint | no | Epoch ns — the source’s time, not ours |
Indexes. aiflow_trigger_firing_tenant_fired answers “what has fired here lately”;
aiflow_trigger_firing_tenant_trigger answers “everything this trigger has done”;
aiflow_trigger_firing_run is the reverse walk from a run.
The refused rows are why this table exists. A webhook refused for an unknown run starts
nothing and writes no directory row, so without this it leaves no trace at all — and those are
exactly the firings somebody asks about later.
action, definition_id and event are recorded even on a refusal, because “a webhook fired and
was supposed to start deploy” is the useful half of that record.
Not audited — the row is the audit record. No access block — closed.
Orchestration
Section titled “Orchestration”aiflow_work_item
Section titled “aiflow_work_item”A queued unit of work: what a run needs done when no agent is immediately free. Declared by lib-cluster, and the queue every engine claims from.
| Field | Type | Required | Description |
|---|---|---|---|
id | ulid | yes | Primary key. Final |
tenant_key | string | yes | Claims are per tenant |
engine_id | string | yes | The orchestrator instance |
engine_kind | string | yes | The join key. A claim filters on it, so work is never claimed by the wrong engine |
parent_work_item_id | string | no | Set for a fan-out child |
run_id / task_id / task_name | string | — | What the work is for |
worker_params | object | no | Selector hints |
payload | object | no | The work itself |
status | string | yes | pending, dispatched, completed, failed, cancelled |
priority | int | no | Lower first |
agent_id | string | no | The agent name holding it |
dispatched_at / completed_at | bigint | no | Epoch ns |
lease_expires_at | bigint | no | When an unfinished claim is swept |
leased_by | string | no | Who holds the lease |
next_attempt_at | bigint | no | Backoff |
retry_count / max_retries | int | no | Attempts |
error | string | no | Last failure |
Indexes. aiflow_work_item_status_priority_idx serves the claim; aiflow_work_item_lease_idx
serves the expiry sweep; aiflow_work_item_run_idx and aiflow_work_item_parent_idx serve the
walks.
The claim is an atomic compare-and-set at the database, not a read-then-write. That is what lets several orchestrators share one queue without handing the same item to two agents.
Not audited. No access block — closed.
aiflow_affinity_binding
Section titled “aiflow_affinity_binding”A key pinned to an agent, so work that must land in one place does. Declared by lib-cluster.
| Field | Type | Required | Description |
|---|---|---|---|
id | ulid | yes | Primary key. Final |
affinity_key | string | yes | Unique. The scope — a run, a definition, a payload key |
agent_id | string | yes | The agent name it is pinned to |
The unique constraint is the mechanism: two orchestrators racing to pin the same key means one insert wins and the other reads the winner, rather than both proceeding with different agents.
Not audited. No access block — closed.
aiflow_agent
Section titled “aiflow_agent”The registry of remote workers. Declared by lib-cluster.
| Field | Type | Required | Description |
|---|---|---|---|
id | ulid | yes | Primary key. Final |
name | string | yes | Unique. The logical agent id every other table stores |
http_address / grpc_address / infra_address | string | no | Where to reach it |
agentversion | string | no | Reported at registration |
binarypath | string | no | |
agenttype | string | no | remote, local |
status | string | no | active, draining, inactive |
metrics | object | no | CPU, memory, GPU and disk from the heartbeat |
metadata | object | no | Tags, loaded models, capacity, tenant list |
Index. aiflow_agent_name_key, unique on name — a duplicate would make create-or-update act on
whichever row came back first, while the other went on advertising stale state.
Audited. No access block — closed; agents are read through /agents.
aiflow_engine
Section titled “aiflow_engine”That a per-tenant engine was started, and its liveness. Declared by lib-cluster.
| Field | Type | Required | Description |
|---|---|---|---|
id | ulid | yes | Primary key. Final |
tenantkey | string | yes | |
engineid | string | yes | The orchestrator instance |
status | string | no |
Index. aiflow_engine_tenant_engine_uq, unique on (tenantkey, engineid).
Not audited. No access block — closed.
Relations
Section titled “Relations”| Name | Parent | Child | Type |
|---|---|---|---|
run_task | aiflow_run.id | aiflow_task.runid | one-to-many |
run_checkpoint | aiflow_run.id | aiflow_checkpoint.run_id | one-to-many |
agent_task | aiflow_agent.name | aiflow_task.workerid | one-to-many |
agent_work_item | aiflow_agent.name | aiflow_work_item.agent_id | one-to-many |
agent_affinity_binding | aiflow_agent.name | aiflow_affinity_binding.agent_id | one-to-many |
The three agent relations join on name, not id. That is neither a typo nor an oversight —
it is verified against live rows.
aiflow_engine_run declares no relation to the engine run tables, and could not usefully: its
run_id matches aiflow_run.id for a DAG run, aiflow_graph_run.run_id for a graph and
aiflow_machine_run.run_id for a machine — which one depends on engine_kind. Resolve it in that
order, or call GET /runs/:id and let the service do it.
Tracing a run back to what caused it
Section titled “Tracing a run back to what caused it”SELECT r.run_id, r.engine_kind, r.definition_id, f.outcome, f.payloadFROM aiflow_engine_run rJOIN aiflow_trigger_firing f ON f.id = r.trigger_firing_idWHERE r.triggered_by = 'on-push'ORDER BY r.createdon DESC;And the firings that caused nothing, which have no other trace:
SELECT trigger_name, outcome, error, fired_onFROM aiflow_trigger_firingWHERE outcome IN ('refused', 'failed')ORDER BY fired_on DESCLIMIT 50;