Skip to content
Talk to our solutions team

Jobs API

GET /jobs jobs defined for this product
GET /runs run history
GET /agents connected agents and their state
GET /workflow/queue queued work
GET /workflow/queue/{id} one queued item

The platform tenancy headers apply, X-Kis-Tenant, X-Kis-Product, X-Kis-Environment.

GET /runs returns run history. A run carries its outcome and its logs, which is what makes “why did last night’s job fail” a query rather than an archaeology exercise across agent hosts.

?id= selects one run, ?filter= takes an expression, ?page= and ?pagesize= paginate. Newest first.

GET /runs/:id returns one run’s state:

{
"runid": "01HQ3M4N5P",
"job": "nightly-reconcile",
"status": "inprogress",
"terminal": false,
"attempt": 1,
"maxattempts": 2,
"concurrency": "forbid",
"progress": { "percent": 40, "message": "page 3 of 7" },
"timing": { "queuedon": "...", "claimedon": "...", "startedon": "...", "durationms": 8120 },
"heldby": "jobs-7d9f",
"lasterror": ""
}

durationms is finishedon − startedon, so it is how long the job ran rather than how long it existed.

GET /runs/:id/logs returns what the job said, oldest first.

GET /runs/01HQ3M4N5P/logs?after=42&limit=200

after is a sequence number. Pass the last one you saw and you get only what is new, which is what makes following a running job cheap.

{ "runid": "01HQ3M4N5P",
"entries": [
{ "seq": 43, "at": "...", "level": "info", "message": "imported", "data": { "rows": 412 } }
] }

These are the job’s own logs. The platform’s operational logs are a separate stream and do not appear here.

POST /runs/:id/retry re-runs a failed run. It re-dispatches the same job with the same input rather than re-deriving it, which matters when the trigger was a one-off webhook you cannot replay. It returns a new run id: the failed run stays as the record of what happened.

{ "message": "job retry submitted", "runid": "01HQ3M4N5Q", "retriedfrom": "01HQ3M4N5P" }

When the orchestrator hands work to an agent, the payload is:

FieldNotes
JobNameThe job being run
RunIDThis execution
LanguageScript language for the job body
CodeThe job body
EnvironmentEnvironment map available to the run
CapabilitiesWhat this job is permitted to do
ConstraintsPlacement constraints for agent selection
ContextTrigger context. The webhook body, the changed row, the commit
NamespacesScript namespaces the job may use; unset means defaults

Namespaces is the field worth being deliberate about. It is the sandbox boundary for the job body, the same way it is for any script, a job that only needs http should not be handed vault.

Context is how trigger data reaches the job: a webhook trigger puts the request body here, a database trigger the changed row, a git-poll trigger the commit.

FieldNotes
JobName, RunIDIdentify the run
StatusOutcome
LogsCaptured output

POST /runs/:id/cancel asks a run to stop.

{ "runid": "01HQ3M4N5P", "status": "inprogress",
"cancelled": true, "immediate": false,
"message": "cancel requested; the run will stop shortly" }

A run that has not started yet is cancelled outright (immediate: true). One already executing is interrupted — the request is recorded and whichever instance holds the run acts on it, so it is not instantaneous. A run that has already finished is left alone and the response says so.

cancelled is a terminal status of its own and is not counted as a failure. A cancelled run is never retried.

Some work cannot be stopped mid-flight — a script runtime that does not support interruption, or a job blocked in a network call. The run is still recorded cancelled, and its lasterror says the work may still be running rather than pretending otherwise.

GET /queue reports what is waiting and what the instance answering is carrying.

{ "tenant": "...",
"waiting": { "created": 42, "claimed": 3, "inprogress": 7 },
"instance": { "instance": "jobs-7d9f", "inflight": 7, "capacity": 100, "free": 25 } }

waiting comes from the database, so every instance gives the same answer. instance is local, so two instances answer differently on purpose — the sum across them is the fleet’s in-flight work.

GET /agents lists connected agents and their state. Throughput is a function of how many are connected, so a silent drop here is capacity you believe you have and do not, see Operations.

The platform’s structured error body: a stable code, a message, and the request id.

StatusMeaning
400Malformed request
401 / 403Authentication or authorisation failure
404Unknown job or run for this tenant
5xxDispatch failure

Every route requires a bearer token:

Authorization: Bearer <token>

The health and readiness probes are the only unauthenticated surface.

The four CEPT headers, X-Customer, X-Product, X-Env, X-Tenant, scope every job, run and queue entry, a request that cannot resolve a tenant is rejected rather than served against a default.

StatusMeaning
400Request parse or validation failure
401Missing or invalid token, as plain text
404No such job, run or queue entry
500Unexpected server error

400 and 500 carry the standard envelope; 401 does not.