Jobs Operations
Growing from one process
Section titled “Growing from one process”A single job server needs a database and nothing else. It runs its own triggers, its own queue and its own jobs. There is no orchestrator to configure and no agent to start.
Growth is a dial rather than a rewrite, and the same job definitions apply at every step:
| one process | runs everything itself |
| several job servers | the database hands each run to exactly one of them, by spare capacity |
| job servers plus agents | work moves to agents as soon as one connects |
| a dedicated fleet | job servers supervise, agents execute |
Running more than one job server needs jobs.scheduler.distributed: true, which is what stops
every instance firing the same scheduled job. Without it each instance runs its own copy of
every schedule. The service warns at startup when it is not set.
Capacity
Section titled “Capacity”Throughput is a function of connected agents. The queue absorbs bursts; sustained queue growth means the trigger rate exceeds what the fleet can drain.
Scale by running more agents. There is no per-job concurrency setting to tune first, that is deliberate, since the usual cause of a backed-up queue is not enough workers rather than badly-tuned ones.
Each job server takes only as much work as it can supervise, so adding one spreads the load rather than duplicating it.
Failure modes
Section titled “Failure modes”An agent disconnects mid-run. Its runs are returned to the queue for another agent, and the attempt is given back rather than charged against the job’s retries: the machine going away is not the job failing. Runs are not lost on agent restart.
A job server dies mid-run. Each claimed run carries a lease the owning server renews while it works. When it stops renewing, another server takes the run over. Recovery is within about one lease period rather than immediate.
A job hangs. Every run has a timeout, an hour unless the job says otherwise. A run that exceeds it is failed and not retried, since a job that hung once will hang again. The timeout interrupts the work rather than only recording it, so the capacity comes back.
Two things still cannot be interrupted: a script runtime that does not support it, and a job blocked in a network call. Those keep their agent slot until they return on their own, and the run says so rather than claiming it stopped.
A job fails. The run records the failure with logs. Retry by id re-dispatches with the original input.
The queue grows without draining. Either agents are gone or a job is consuming them. Check agent count first. A silent drop in connected agents is capacity you believe you have and do not.
A trigger stops firing. Scheduled triggers are the ones to watch; a cron that silently stops produces no error, just an absence. Alert on expected runs not happening, not only on runs failing.
Jobs, workflows and flows
Section titled “Jobs, workflows and flows”Three things on the platform run multi-step work, and they are not interchangeable:
| Use | When |
|---|---|
| Jobs | Background work on a schedule or a trigger |
| Workflows | Long-running business processes with human steps |
| AI Flow | DAGs of tasks, including model calls and agentic loops |
The distinction that matters: a job is fired, a workflow waits. If your process needs to pause for a human approval and resume days later, that is a workflow.
What to watch
Section titled “What to watch”| Signal | Why it matters |
|---|---|
| Queue depth trend | Not enough agents for the trigger rate |
| Connected agent count | A silent drop is invisible capacity loss |
| Run failure rate by job | One job degrading, not the fleet |
| Retry rate | Work that only succeeds on a second attempt is still a defect |
| Expected-but-missing runs | A trigger that stopped firing produces no error |