Jobs API
GET /jobs jobs defined for this productGET /runs run historyGET /agents connected agents and their stateGET /workflow/queue queued workGET /workflow/queue/{id} one queued itemRequired headers
Section titled “Required headers”The platform tenancy headers apply, X-Kis-Tenant, X-Kis-Product, X-Kis-Environment.
GET /runs returns run history. A run carries its outcome and its logs, which is what makes
“why did last night’s job fail” a query rather than an archaeology exercise across agent hosts.
?id= selects one run, ?filter= takes an expression, ?page= and ?pagesize= paginate.
Newest first.
GET /runs/:id returns one run’s state:
{ "runid": "01HQ3M4N5P", "job": "nightly-reconcile", "status": "inprogress", "terminal": false, "attempt": 1, "maxattempts": 2, "concurrency": "forbid", "progress": { "percent": 40, "message": "page 3 of 7" }, "timing": { "queuedon": "...", "claimedon": "...", "startedon": "...", "durationms": 8120 }, "heldby": "jobs-7d9f", "lasterror": ""}durationms is finishedon − startedon, so it is how long the job ran rather than how long
it existed.
GET /runs/:id/logs returns what the job said, oldest first.
GET /runs/01HQ3M4N5P/logs?after=42&limit=200after is a sequence number. Pass the last one you saw and you get only what is new, which is
what makes following a running job cheap.
{ "runid": "01HQ3M4N5P", "entries": [ { "seq": 43, "at": "...", "level": "info", "message": "imported", "data": { "rows": 412 } } ] }These are the job’s own logs. The platform’s operational logs are a separate stream and do not appear here.
POST /runs/:id/retry re-runs a failed run. It re-dispatches the same job with the same
input rather than re-deriving it, which matters when the trigger was a one-off webhook you
cannot replay. It returns a new run id: the failed run stays as the record of what
happened.
{ "message": "job retry submitted", "runid": "01HQ3M4N5Q", "retriedfrom": "01HQ3M4N5P" }What a job dispatch carries
Section titled “What a job dispatch carries”When the orchestrator hands work to an agent, the payload is:
| Field | Notes |
|---|---|
JobName | The job being run |
RunID | This execution |
Language | Script language for the job body |
Code | The job body |
Environment | Environment map available to the run |
Capabilities | What this job is permitted to do |
Constraints | Placement constraints for agent selection |
Context | Trigger context. The webhook body, the changed row, the commit |
Namespaces | Script namespaces the job may use; unset means defaults |
Namespaces is the field worth being deliberate about. It is the sandbox boundary for the job
body, the same way it is for any script, a job that only
needs http should not be handed vault.
Context is how trigger data reaches the job: a webhook trigger puts the request body here, a
database trigger the changed row, a git-poll trigger the commit.
The response
Section titled “The response”| Field | Notes |
|---|---|
JobName, RunID | Identify the run |
Status | Outcome |
Logs | Captured output |
POST /runs/:id/cancel asks a run to stop.
{ "runid": "01HQ3M4N5P", "status": "inprogress", "cancelled": true, "immediate": false, "message": "cancel requested; the run will stop shortly" }A run that has not started yet is cancelled outright (immediate: true). One
already executing is interrupted — the request is recorded and whichever instance
holds the run acts on it, so it is not instantaneous. A run that has already finished
is left alone and the response says so.
cancelled is a terminal status of its own and is not counted as a failure. A
cancelled run is never retried.
Some work cannot be stopped mid-flight — a script runtime that does not support
interruption, or a job blocked in a network call. The run is still recorded cancelled,
and its lasterror says the work may still be running rather than pretending
otherwise.
GET /queue reports what is waiting and what the instance answering is carrying.
{ "tenant": "...", "waiting": { "created": 42, "claimed": 3, "inprogress": 7 }, "instance": { "instance": "jobs-7d9f", "inflight": 7, "capacity": 100, "free": 25 } }waiting comes from the database, so every instance gives the same answer.
instance is local, so two instances answer differently on purpose — the sum across
them is the fleet’s in-flight work.
Agents
Section titled “Agents”GET /agents lists connected agents and their state. Throughput is a function of how many are
connected, so a silent drop here is capacity you believe you have and do not, see
Operations.
Errors
Section titled “Errors”The platform’s structured error body: a stable code, a message, and the request id.
| Status | Meaning |
|---|---|
400 | Malformed request |
401 / 403 | Authentication or authorisation failure |
404 | Unknown job or run for this tenant |
5xx | Dispatch failure |
See also
Section titled “See also”- Usage: triggers, runs and retries
- Operations: capacity and failure modes
Authentication
Section titled “Authentication”Every route requires a bearer token:
Authorization: Bearer <token>The health and readiness probes are the only unauthenticated surface.
Tenancy
Section titled “Tenancy”The four CEPT headers, X-Customer, X-Product, X-Env, X-Tenant, scope every job, run and
queue entry, a request that cannot resolve a tenant is rejected rather than served against a
default.
Errors
Section titled “Errors”| Status | Meaning |
|---|---|
400 | Request parse or validation failure |
401 | Missing or invalid token, as plain text |
404 | No such job, run or queue entry |
500 | Unexpected server error |
400 and 500 carry the standard envelope;
401 does not.