Skip to content
Talk to our solutions team

Entity reference

AI Flow persists its state as entities on the platform’s Data engine, in the tenant’s own schema. They are ordinary entities: the same access rules, row-level security, compliance directives and audit stream apply, and they are reachable through the generic entity API as well as through this service’s own routes.

Eleven entities, all prefixed aiflow_.

This is the organising idea, and the reason there is no single “runs” table.

A DAG run is a set of nodes with dependencies plus a row per node attempt. A graph run is a position and a step count. A state-machine run is a state name and a context map. Those are genuinely different shapes, and one table holding all three would be a union of columns where most are null for any given row.

So each engine persists what it has, and a directory sits over them:

EngineWritesHolds
dagaiflow_run, aiflow_taskA run, and one row per node attempt
graphaiflow_graph_runCurrent node, step count, the path walked
statemachineaiflow_machine_runCurrent state, context, history
datapipenot herePipes belong to datapipe.svc, under its own datapipeline_ prefix
all of themaiflow_engine_runThe directory: which engine owns which run

aiflow_engine_run is what makes “I have a run id — which engine ran it?” answerable. It holds no state of its own, so it cannot disagree with an engine about one, and the engine that owns a run writes its directory row on the same save that writes its own state. One writer, no drift.

The tables are not all AI Flow’s. Two libraries beneath it contribute entity fragments and the service composes them at boot; AI Flow itself declares none.

That matters when you are reading a schema and wondering where a column’s meaning is defined, and it is why every service that registers engines gets the same directory and queue tables.

Declared byEntitiesWhy it lives there
lib-cluster (the substrate)aiflow_engine_run, aiflow_trigger_firing, aiflow_engine, aiflow_agent, aiflow_work_item, aiflow_affinity_bindingEvery service registering engines needs a run directory, a firing log, a work queue and an agent registry. Declaring them once is why they agree across services.
lib-guppi (the engines)aiflow_run, aiflow_task, aiflow_graph_run, aiflow_machine_run, aiflow_checkpointEngine-specific run state — each engine’s own shape.
AI Flow—Contributes no fragment. Everything it stores belongs to a layer below.

The aiflow_ prefix is applied by the datastore rather than baked into the fragments, which is how four services compose the same tables under four different prefixes.

aiflow_engine_run ── the directory over every engine's runs
│
├── dag ─────────── aiflow_run ──< aiflow_task
│ a flow run one node attempt
│ └──< aiflow_checkpoint
│ iterator progress
├── graph ───────── aiflow_graph_run
└── statemachine ── aiflow_machine_run
aiflow_trigger_firing what a trigger sent, and what it did
aiflow_work_item a queued task >── aiflow_agent
aiflow_affinity_binding a key pinned >── a registered worker
aiflow_engine an engine's liveness

The agent relations join on name, not id. aiflow_agent has both — id is a generated ULID, name is the logical agent id (harness-agent-1) — and every other table stores the name: aiflow_work_item.agent_id, aiflow_affinity_binding.agent_id and aiflow_task.workerid all hold harness-agent-1, never the ULID.

All eleven inherit kisai.common, so these six exist on every table without being declared:

ColumnTypeNullableSet by
createdbystringnoThe caller’s identity, on create. Final
createdontimestampnoCreation time. Final
updatedbystringyesThe caller’s identity, on each write
updatedontimestampyesLast write time
deletedbystringyesSoft delete
deletedontimestampyesSoft delete

createdby is load-bearing rather than decorative on the user-facing tables: reads are scoped to createdby = <caller>, built structurally by the handler rather than written as a rule.

Ordering. Every entity a person lists is createdon desc, id desc, so a list is newest-first. The exceptions order by what they are dispatched or addressed by: aiflow_agent by name, aiflow_work_item by priority, createdon, aiflow_affinity_binding by createdon, aiflow_checkpoint by run_id, node_name, and aiflow_trigger_firing by fired_on desc.

Timestamps are of two kinds, deliberately. createdon and its siblings are real timestamptz columns from kisai.common. Columns an engine writes itself — starttime, endtime, ended_on, fired_on, lease_expires_at — are bigint epoch nanoseconds, because they are compared and arithmetic’d on the hot path.

Reading one as the other is the common mistake, and it fails quietly: a timestamptz read as an integer yields 0, which renders as year one rather than as an error.

Audit. An entity records nothing until its audit: block says so. Where it is on, field names and primary keys are recorded, never values. On this service that is aiflow_run, aiflow_task and aiflow_agent. See Operations.

Access. The engine denies by default, so an operation with no rule is refused for every external caller. Each entity opts in only to what this service’s own handlers perform, and the service writes execution state under its own service context, which is not gated.

Eight entities carry no access block at all, which IS the closed posture: aiflow_engine_run, aiflow_trigger_firing, aiflow_engine, aiflow_agent, aiflow_work_item, aiflow_affinity_binding, aiflow_checkpoint, and the two native-engine run tables aiflow_graph_run and aiflow_machine_run. They are the orchestrator’s own bookkeeping; a caller reaches what they hold through /runs, /triggers and the agent routes rather than by reading rows.

The actions block is sealed (access-lock) throughout. A product layer composed on top may not replace it and grant itself operations this service refused. It may still add rls or fields rules, which can only narrow what a caller reaches.


Which engine owns which run. One row per run, whatever engine ran it, and no state — asked for a run’s detail, the service asks the engine that owns it.

Declared by lib-cluster. Written by the engine on the same save as its own run row; a trigger fills in cause. Recording merges rather than replaces, so whichever fact arrives second completes the row instead of erasing the other’s half.

FieldTypeRequiredDescription
idulidyesPrimary key. Final
run_idstringyesThe engine’s own run id
engine_kindstringyesdag, graph, statemachine, datapipe, test
tenant_keystringyesThe full tenant key
definition_idstringnoWhat was run
parent_run_idstringnoSet when another run spawned this one
triggered_bystringnoWhich trigger caused it, by name. Empty when a caller started it
trigger_firing_idstringnoWhich firing — and therefore what arrived
ended_onbigintnoEpoch ns. Null while live; first write wins

triggered_by and trigger_firing_id are not redundant. The first says a trigger caused this run; the second says which firing, so the payload is reachable. A trigger that fires hourly leaves a dozen candidate causes a day, and matching them by timestamp is a guess.

Indexes. aiflow_engine_run_tenant_run_uq, unique on (tenant, run), which is what makes recording idempotent — an engine writes the directory on every save, so a create-only write would fail on the second step of every run. aiflow_engine_run_tenant_created serves the newest-first listing.

The row stays after a run ends. “Which engine ran this” is asked about finished runs too, by exactly the operator the directory is for.


One execution of a DAG definition. Declared by lib-guppi.

FieldTypeRequiredDescription
idulidyesPrimary key. Final
definitionidstringyesThe definition executed
runtypestringyesdag. Historical rows may hold datapipeline
contextobjectnoThe run’s shared data structure
metadataobjectnoWorker assignments, tasklist, affinity key, the execution plan
statesobjectnoState history; the last entry is the current state
starttimebigintnoEpoch ns
endtimebigintnoEpoch ns
durationbigintnoendtime - starttime

Indexes. aiflow_run_owner_recent_idx on (createdby, createdon) — the ownership predicate is unconditional on every user-facing read; aiflow_run_definition_idx on definitionid.

Audited. Access: actions sealed.

One node-execution attempt within a run, bound to a worker. Declared by lib-guppi.

FieldTypeRequiredDescription
idulidyesPrimary key. Final
namestringyesThe node name. A fan-out attempt carries a synthetic suffix
definitionidstringyesDenormalised, for the definition index
runidstringyesParent run — the run_task relation
workeridstringnoThe agent name that executed it, not a ULID
contextobjectnoThe task’s own data
metadataobjectnoParallel group, pending events
statesobjectnoState history
starttime / endtime / durationbigintnoEpoch ns
resumedataobjectnoWhat a suspended node needs in order to continue

Indexes. aiflow_task_run_node_idx on (runid, name); aiflow_task_definition_idx.

Audited. Access: actions sealed.

A graph may cycle, so a run is a position plus a bound. Declared by lib-guppi.

FieldTypeRequiredDescription
idulidyesPrimary key. Final
run_idstringyesUnique. The engine’s own id
definition_idstringyesWhat is being walked
taskstringnoThe current node
stepsbigintnoSteps taken, checked against the definition’s maxsteps
pathtextnoThe nodes walked, in order
run_countstextnoPer-node visit counts — a cycle’s bookkeeping
statetextnoThe run’s own data
statusstringnoEngine status
reasonstringnoWhy it stopped
errortextnoFailure detail

Index. aiflow_graph_run_runid_uq, unique on run_id.

Not audited. No access block — closed.

aiflow_machine_run — the state-machine engine

Section titled “aiflow_machine_run — the state-machine engine”

A machine rests between events, so a run is a state and what it carries. Declared by lib-guppi.

FieldTypeRequiredDescription
idulidyesPrimary key. Final
run_idstringyesUnique
definition_idstringyesThe machine
statestringnoThe state it is resting in
contexttextnoThe machine’s data
historytextnoStates visited, in order
liveboolnoFalse once terminal
suspendedboolnoWaiting on something other than an event
failuretextnoFailure detail
reasonstringnoWhy it stopped

Index. aiflow_machine_run_runid_uq, unique on run_id.

Not audited. No access block — closed.

Iterator progress for a node that fans out, so a resumed run does not repeat work. Declared by lib-guppi; run_checkpoint is the relation to aiflow_run.

FieldTypeRequiredDescription
idulidyesPrimary key. Final
run_idstringyesThe run
node_namestringyesThe fanning node
definition_idstringnoDenormalised
iterator_typestringnomap, foreach
cursortextnoWhere to resume
idx / totalbigintnoPosition and size
doneboolnoIteration complete
updated_atbigintnoEpoch ns

Index. aiflow_checkpoint_run_node_uq, unique on (run_id, node_name) — one checkpoint per fanning node, so a second write updates rather than duplicating.

Not audited. No access block — closed.


What an inbound trigger sent, and what became of it. Declared by lib-cluster.

Append-only. Unlike the directory, which an engine rewrites as a run advances, a firing is a fact about one instant: written once, after the outcome is known. There is no update path and no delete path, because an audit trail that can be edited is not one.

FieldTypeRequiredDescription
idulidyesPrimary key. Final
tenant_keystringyesEvery read is qualified by it — an id alone is not a key
trigger_namestringyesWhich trigger fired
trigger_typestringnowebhook, cron, gitpoll, filesystem
payloadtextnoWhat arrived, as JSON. The whole reason the table exists
actionstringnostart or signal — what the binding meant to do
definition_idstringnoFor a start
eventstringnoFor a signal
run_idstringnoWhat it started or moved. Empty on a refusal
outcomestringnostarted, signalled, refused, failed
errortextnoWhy, on a refusal or a failure
fired_onbigintnoEpoch ns — the source’s time, not ours

Indexes. aiflow_trigger_firing_tenant_fired answers “what has fired here lately”; aiflow_trigger_firing_tenant_trigger answers “everything this trigger has done”; aiflow_trigger_firing_run is the reverse walk from a run.

The refused rows are why this table exists. A webhook refused for an unknown run starts nothing and writes no directory row, so without this it leaves no trace at all — and those are exactly the firings somebody asks about later.

action, definition_id and event are recorded even on a refusal, because “a webhook fired and was supposed to start deploy” is the useful half of that record.

Not audited — the row is the audit record. No access block — closed.


A queued unit of work: what a run needs done when no agent is immediately free. Declared by lib-cluster, and the queue every engine claims from.

FieldTypeRequiredDescription
idulidyesPrimary key. Final
tenant_keystringyesClaims are per tenant
engine_idstringyesThe orchestrator instance
engine_kindstringyesThe join key. A claim filters on it, so work is never claimed by the wrong engine
parent_work_item_idstringnoSet for a fan-out child
run_id / task_id / task_namestring—What the work is for
worker_paramsobjectnoSelector hints
payloadobjectnoThe work itself
statusstringyespending, dispatched, completed, failed, cancelled
priorityintnoLower first
agent_idstringnoThe agent name holding it
dispatched_at / completed_atbigintnoEpoch ns
lease_expires_atbigintnoWhen an unfinished claim is swept
leased_bystringnoWho holds the lease
next_attempt_atbigintnoBackoff
retry_count / max_retriesintnoAttempts
errorstringnoLast failure

Indexes. aiflow_work_item_status_priority_idx serves the claim; aiflow_work_item_lease_idx serves the expiry sweep; aiflow_work_item_run_idx and aiflow_work_item_parent_idx serve the walks.

The claim is an atomic compare-and-set at the database, not a read-then-write. That is what lets several orchestrators share one queue without handing the same item to two agents.

Not audited. No access block — closed.

A key pinned to an agent, so work that must land in one place does. Declared by lib-cluster.

FieldTypeRequiredDescription
idulidyesPrimary key. Final
affinity_keystringyesUnique. The scope — a run, a definition, a payload key
agent_idstringyesThe agent name it is pinned to

The unique constraint is the mechanism: two orchestrators racing to pin the same key means one insert wins and the other reads the winner, rather than both proceeding with different agents.

Not audited. No access block — closed.

The registry of remote workers. Declared by lib-cluster.

FieldTypeRequiredDescription
idulidyesPrimary key. Final
namestringyesUnique. The logical agent id every other table stores
http_address / grpc_address / infra_addressstringnoWhere to reach it
agentversionstringnoReported at registration
binarypathstringno
agenttypestringnoremote, local
statusstringnoactive, draining, inactive
metricsobjectnoCPU, memory, GPU and disk from the heartbeat
metadataobjectnoTags, loaded models, capacity, tenant list

Index. aiflow_agent_name_key, unique on name — a duplicate would make create-or-update act on whichever row came back first, while the other went on advertising stale state.

Audited. No access block — closed; agents are read through /agents.

That a per-tenant engine was started, and its liveness. Declared by lib-cluster.

FieldTypeRequiredDescription
idulidyesPrimary key. Final
tenantkeystringyes
engineidstringyesThe orchestrator instance
statusstringno

Index. aiflow_engine_tenant_engine_uq, unique on (tenantkey, engineid).

Not audited. No access block — closed.


NameParentChildType
run_taskaiflow_run.idaiflow_task.runidone-to-many
run_checkpointaiflow_run.idaiflow_checkpoint.run_idone-to-many
agent_taskaiflow_agent.nameaiflow_task.workeridone-to-many
agent_work_itemaiflow_agent.nameaiflow_work_item.agent_idone-to-many
agent_affinity_bindingaiflow_agent.nameaiflow_affinity_binding.agent_idone-to-many

The three agent relations join on name, not id. That is neither a typo nor an oversight — it is verified against live rows.

aiflow_engine_run declares no relation to the engine run tables, and could not usefully: its run_id matches aiflow_run.id for a DAG run, aiflow_graph_run.run_id for a graph and aiflow_machine_run.run_id for a machine — which one depends on engine_kind. Resolve it in that order, or call GET /runs/:id and let the service do it.

SELECT r.run_id, r.engine_kind, r.definition_id, f.outcome, f.payload
FROM aiflow_engine_run r
JOIN aiflow_trigger_firing f ON f.id = r.trigger_firing_id
WHERE r.triggered_by = 'on-push'
ORDER BY r.createdon DESC;

And the firings that caused nothing, which have no other trace:

SELECT trigger_name, outcome, error, fired_on
FROM aiflow_trigger_firing
WHERE outcome IN ('refused', 'failed')
ORDER BY fired_on DESC
LIMIT 50;