Skip to content
Talk to our solutions team

Jobs Usage

GET /jobs jobs defined for the product, and what is scheduled
GET /runs run history
GET /runs/:id one run: status, progress, timing, why it stopped
GET /runs/:id/logs what the job said, in order
POST /runs/:id/cancel ask a running job to stop
POST /runs/:id/retry re-run a failed run
GET /agents connected agents and their state
GET /workflow/queue queued work
GET /workflow/queue/:id one queued item

Runs carry their outcome and their logs, so “why did last night’s job fail” is a query rather than an archaeology exercise across agent hosts.

A job runs when something fires it.

A cron expression, the common case, and the one to reach for when the work is genuinely periodic rather than reactive.

An inbound HTTP call fires the job, with the request body as input. Use this for third-party callbacks, payment confirmations, delivery notifications, where the other side pushes.

A change in a watched table fires the job. This removes the polling loop most teams write by hand: a job that should run “when a row lands” does not need a one-minute cron and a watermark column, with the lag and the missed-row edge cases that come with them.

A new commit on a watched repository fires the job, the basis for content or configuration pipelines that follow a repo.

A message in a connected workspace fires the job. The “run this from chat” pattern, without building a bot.

Beside triggers:, four optional fields decide what happens when things do not go smoothly. All four default to the behaviour jobs had before they existed, so an existing job is unaffected by leaving them out.

jobs:
- name: nightly-reconcile
retries: 2 # default 0, a failure is final
concurrency: forbid # default allow
timeout: 30m # default 1h
agent:
tag: gpu
affinity: sticky

retries re-runs a failed run, with a growing gap between attempts. The default is 0 because the platform cannot know whether your job is safe to run twice, only you can.

concurrency answers whether the job may run while a run of it is already going.

allowany number at once. Right for work that handles its own slice, a row, a message
queuewaits for the active run, then goes. Nothing is lost, nothing overlaps
forbidthe new run is marked skipped. Right for a poll or a refresh, where the run that was due while the last one ran has been superseded by the next one anyway

timeout bounds one run, measured from when it actually started rather than when it was queued, so a run that waited for capacity is not charged for waiting. A run that exceeds it is failed and not retried, even with retries left: a job that hung once will hang again, and retrying turns one stuck slot into a slow loop of them.

The timeout stops the work, not just the record — a runaway loop is interrupted rather than left running. Work that cannot be interrupted, such as a job blocked in a network call, is reported as such rather than silently assumed stopped.

agent decides where it runs. tag matches an agent label, affinity: sticky prefers wherever the job ran last, which is worth having when there is a warm cache or a checked-out repository to reuse. Sticky is a preference, so a job whose usual agent is gone still runs.

A job can say what it is doing while it does it.

progress(40, "page 3 of 7")
logdata("imported", { rows: 412, skipped: 3 })
pluginlog("done")

These appear on the run as it goes, not when it finishes, so a long job is watchable rather than opaque, and a run whose agent dies still has everything said up to that point.

These are your logs, kept with the run and queryable by id. They are separate from the platform’s own operational logs, which is deliberate, you should not have to read past connection-pool statistics to find out how many invoices your job imported.

GET /runs/:id/logs?after=<seq> returns only what is new, so following a running job does not mean re-reading its whole log each time.

Every firing produces a run with a status, timing, output and logs. A failed run can be retried by id.

Retry re-dispatches the same job with the same input rather than re-deriving it. That matters when the trigger was a one-off webhook you cannot replay, re-deriving would mean the input is simply gone.

A run moves through created → claimed → inprogress and ends completed, failed or skipped. skipped is not a failure, it means a run of that job was already going and the job declares concurrency: forbid.

Jobs execute on agents: long-lived worker processes that connect and receive work. The orchestrator distributes runs across whatever is connected, so scaling throughput is running more agents rather than reconfiguring anything.

An agent that disconnects mid-run leaves the run recoverable rather than lost: its work is returned to the queue for another agent, and the attempt is given back rather than charged against the job’s retries, since the machine going away is not the job failing.

You do not need agents to start. A single job server runs its own work, and moves to agents the moment one connects, without changing anything about the jobs.

  • Operations: capacity, failure modes and what to watch