Skip to content
Talk to our solutions team

Jobs Operations

A single job server needs a database and nothing else. It runs its own triggers, its own queue and its own jobs. There is no orchestrator to configure and no agent to start.

Growth is a dial rather than a rewrite, and the same job definitions apply at every step:

one processruns everything itself
several job serversthe database hands each run to exactly one of them, by spare capacity
job servers plus agentswork moves to agents as soon as one connects
a dedicated fleetjob servers supervise, agents execute

Running more than one job server needs jobs.scheduler.distributed: true, which is what stops every instance firing the same scheduled job. Without it each instance runs its own copy of every schedule. The service warns at startup when it is not set.

Throughput is a function of connected agents. The queue absorbs bursts; sustained queue growth means the trigger rate exceeds what the fleet can drain.

Scale by running more agents. There is no per-job concurrency setting to tune first, that is deliberate, since the usual cause of a backed-up queue is not enough workers rather than badly-tuned ones.

Each job server takes only as much work as it can supervise, so adding one spreads the load rather than duplicating it.

An agent disconnects mid-run. Its runs are returned to the queue for another agent, and the attempt is given back rather than charged against the job’s retries: the machine going away is not the job failing. Runs are not lost on agent restart.

A job server dies mid-run. Each claimed run carries a lease the owning server renews while it works. When it stops renewing, another server takes the run over. Recovery is within about one lease period rather than immediate.

A job hangs. Every run has a timeout, an hour unless the job says otherwise. A run that exceeds it is failed and not retried, since a job that hung once will hang again. The timeout interrupts the work rather than only recording it, so the capacity comes back.

Two things still cannot be interrupted: a script runtime that does not support it, and a job blocked in a network call. Those keep their agent slot until they return on their own, and the run says so rather than claiming it stopped.

A job fails. The run records the failure with logs. Retry by id re-dispatches with the original input.

The queue grows without draining. Either agents are gone or a job is consuming them. Check agent count first. A silent drop in connected agents is capacity you believe you have and do not.

A trigger stops firing. Scheduled triggers are the ones to watch; a cron that silently stops produces no error, just an absence. Alert on expected runs not happening, not only on runs failing.

Three things on the platform run multi-step work, and they are not interchangeable:

UseWhen
JobsBackground work on a schedule or a trigger
WorkflowsLong-running business processes with human steps
AI FlowDAGs of tasks, including model calls and agentic loops

The distinction that matters: a job is fired, a workflow waits. If your process needs to pause for a human approval and resume days later, that is a workflow.

SignalWhy it matters
Queue depth trendNot enough agents for the trigger rate
Connected agent countA silent drop is invisible capacity loss
Run failure rate by jobOne job degrading, not the fleet
Retry rateWork that only succeeds on a second attempt is still a defect
Expected-but-missing runsA trigger that stopped firing produces no error