Server Modes
For horizontal scale, run Testing as a control plane (orchestrator) plus workers (agents).
The orchestrator spreads two kinds of run across its agents: a functional run of your suites
(kis test run --orchestrator) and a load run (kis test load --orchestrator).
gRPC (agents dial out) ┌─────────┐ register + heartbeat + work ┌──────────────┐ │ agent │ ───────────────────────────────▶ │ orchestrator │◀── HTTPS API (JWT) └─────────┘ ◀─────────────────────────────── └──────────────┘ runs its share dispatches, results returns results of each run to the submitter- The orchestrator runs a gRPC server (agents register and heartbeat) and an HTTPS API for submitting runs and reading results.
- Agents dial the orchestrator as gRPC clients (only the orchestrator needs an inbound
gRPC port), run the work they are handed, and send results back over the same channel. Each
agent also serves a minimal local HTTPS endpoint:
GET /ready,GET /health.
Both commands read a .test.yaml config file (-f/--config) plus the shared infra flags
(--port, --sid, --config_url, service-discovery and TLS cert flags; see
CLI Reference).
Orchestrator
Section titled “Orchestrator”test.svc orchestrator -f orchestrator.test.yamlKey config:
| Key | Required | Meaning |
|---|---|---|
grpc | yes | gRPC server port (startup fails without it) |
host / port | HTTPS API bind address | |
timeouts.read / timeouts.write / timeouts.readheader | server timeouts (seconds or duration strings; defaults 15s/15s/5s) | |
orchestrator.queue.maxdepth | work-queue depth | |
serviceid | service identity | |
loadtenants | tenants to warm-load at boot | |
grpccreds | gRPC credentials | |
localagent | optional in-process agent: id, serverurl, grpc, heartbeatduration (default 15s), infrahost, grpcaddress, httpaddress |
TLS comes from the container-registered service certificates (--svccert/--svckey).
Shutdown is graceful: SIGINT/SIGTERM drains agent and orchestrator work with 30s timeouts.
HTTP API
Section titled “HTTP API”All routes are JWT-secured.
| Route | Method | Purpose |
|---|---|---|
/execute/run | POST | start a functional run of submitted suites across the registered agents |
/execute/run/:id | GET | the functional run’s status, and each shard’s results when it is done |
/execute/load | POST | start a load run of a resolved profile across the registered agents |
/execute/load/:id | GET | the load run’s status, and the merged result when it is done |
/list/tests | GET | teststeps/testcases/testscenarios/testplans known for the tenant (hierarchy format) |
/execute/tests | POST | fan a filtered hierarchy-format test set out to agents |
/list/test/runs | GET | hierarchy-format run history with metrics and logs |
/agents | GET | registered agents |
/workflow/queue | GET | queue status |
/workflow/queue/:id | DELETE | cancel queued work |
/ready, /health | GET | probes |
Runs are submitted, then polled
Section titled “Runs are submitted, then polled”POST /execute/run and POST /execute/load answer 202 with the run’s id as soon as the run
starts:
{"run_id": "6f1c…", "status": "running"}GET /execute/run/:id (or /execute/load/:id) answers with the run. status is running
until the run ends, then done with its result, or failed with an error saying why the
run could not happen. A finished run stays available for an hour. The clients ask every two
seconds, so a run lasts as long as it needs to while each request stays short.
Functional runs
Section titled “Functional runs”kis test run -t tests/ --orchestrator https://orch.internal:8443 --token "$TALOS_TOKEN"Your machine loads the suites exactly as a local run does, then:
- Splits the run into units. The root suite’s own tests are one unit, or one unit per test
when the root suite is
parallel: true. Each child suite of the root, with everything under it, is one unit. A sequential suite’s tests stay together, because a test may use what an earlier one published. - Runs the root suite’s
before_allhere, once. What it extracted and published goes to the agents with the units. - Packs the suite directory and submits it with the units.
node_modules,vendor,.git,.artifacts,test-resultsandplaywright-reportare left out. When the suites read files outside--tests(a shared fixtures folder, an import from a sibling directory), pass--bundle-rootto send the directory that contains both. - The orchestrator spreads the units across its agents by test count, largest first onto
the least loaded agent.
--max-agents Ncaps how many are used. - Each agent unpacks the directory and loads the suites from it, then runs its units with
the same engine as a local run, starting from what
before_allleft. Every step type runs,playwright:included. - The results come back: everything each unit reported, and the files its steps kept.
Your machine passes them through its own reporters and store, so the console output,
--junit,--html,--dband--artifactsmatch a local run of the same suites. Then it runs the root suite’safter_all, once.
A unit whose agent reports an error, or does not report within --timeout (default 30m), has
each of its tests reported errored with the reason, so the run fails.
A value published with export: run reaches the tests of its own unit. Units may run on
different agents, so publish what every unit needs from the root suite’s before_all.
POST /execute/run body, as kis test run --orchestrator sends it:
{ "run_id": "optional; one is minted if omitted", "bundle": "<the suite directory, a base64 gzipped tar>", "tests": "path of --tests within the bundle", "units": [{"key": ".", "tests": 4}, {"key": "orders", "tests": 12}], "setup": {"scope": {"token": "…"}, "run": {}, "suite": {}}, "tags": ["smoke"], "exclude_tags": [], "max_agents": 0, "timeout_seconds": 1800}When done, result.shards lists each shard: shard, shards, host, the units it ran, its
packed results, and an error when it could not run them.
Load runs
Section titled “Load runs”POST /execute/load body:
{ "run_id": "optional; one is minted if omitted", "profile": { "...": "a resolved load profile, tests already expanded" }, "bundle": "<the suite directory, a base64 gzipped tar>", "bundle_root": "/home/you/project/tests", "max_agents": 0, "grace_seconds": 120}You rarely build this by hand: kis test load --orchestrator <url> resolves the suite and
sends it with the suite directory, packed (bundle, and bundle_root, where the directory is
on your machine). Each agent unpacks the directory and maps the profile’s paths onto its copy,
so the files a step reads are there. Agents need no deployment step when your tests change,
and a suite with a very large table: makes a larger dispatch.
Shards start together. Each agent prepares its shard and reports ready; the orchestrator waits for every dispatched shard, up to two minutes, then starts all that are ready at once. A shard that failed to prepare or was not ready in time is listed with the reason. Each shard also carries a deadline derived from the profile’s own duration, so generators stop when the run should have ended.
When done, result holds the merged result plus one entry per shard:
{ "result": { "Overall": {"...": "..."}, "Thresholds": ["..."] }, "shards": [ {"shard": 1, "shards": 3, "host": "gen-a", "contributed": true, "error": ""}, {"shard": 2, "shards": 3, "host": "gen-b", "contributed": false, "error": "not ready to start within 2m0s"} ]}Every shard is listed whether it contributed or not, and a run that lost a generator exits 1: less load reads as a faster service.
The machine that submits the run runs the suite’s before_all once, before dispatch, and its
after_all once, after the merge; what setup extracted and published travels to the agents with
the profile. Agents run neither, so a setup that seeds a fixed row or claims a unique name runs
once.
Run history
Section titled “Run history”GET /list/test/runs query params: id, filter, page (default 1), pagesize
(default 10, max 100), metricfields, logfields. filter is a condition over the testrun
fields, such as status = 'failed'; one that does not parse is answered 400.
POST /execute/tests body (hierarchy format):
{ "type": "testplan", "pattern": "api-smoke", "vars": {"baseurl": "https://api.staging.example.com"}, "logresponse": false, "workercount": 2}Splits the filtered tests across workercount agents; responds
{"message": "test runs started successfully", "status": {"runid": "..."}}.
test.svc agent -f agent.test.yamlKey config:
| Key | Required | Meaning |
|---|---|---|
serverurl | yes | orchestrator gRPC address to dial |
sid | agent instance id | |
grpc | agent’s own gRPC port | |
heartbeatduration | default 15s | |
infrahost, grpcaddress, httpaddress | addresses advertised on registration | |
tenants | tenant keys this agent serves | |
exclusive | reserve the agent for its tenant list | |
max_concurrent_tasks | concurrency cap | |
host / port + timeouts.* | the local /ready + /health HTTPS server | |
loadtenants, grpccreds | as for the orchestrator |
The agent takes three kinds of work:
- Units of a functional run from
/execute/run: it unpacks the suite directory into a temporary directory, loads the suites, runs its units, and sends back what they reported and the files their steps kept. The directory is removed afterwards. - A load shard from
/execute/load: it unpacks the suite directory, prepares its share of the profile (users and rates divided across the agents, durations not), reports ready, starts when the orchestrator says so, and returns a mergeable snapshot; see Load testing. - A hierarchy-format run from
/execute/tests(run id, environment, tests, tenant): it runs and streams results back. Metric records becometestrunmetricrows, log lines becometestrunlogrows, and a completion or error message closes out thescenariorun.
Browser steps on agents
Section titled “Browser steps on agents”An agent runs playwright: steps with its own Node, Playwright and browser; the suite
directory it receives has no node_modules. Install Node 18 or later and @playwright/test
on the agent, and set NODE_PATH to the global node_modules directory so Node finds
Playwright from any directory. For Obscura, the default browser, put obscura and
obscura-worker on the agent’s PATH (or set TALOS_OBSCURA); for a browser Playwright
launches, install it with npx playwright install:
npm install -g @playwright/testexport NODE_PATH="$(npm root -g)"curl -L https://github.com/h4ckf0r0day/obscura/releases/download/v0.2.3/obscura-x86_64-linux.tar.gz | tar xz -C /usr/local/binnpx playwright install --with-deps chromium # only for runs that choose chromiumThe run’s --browser travels with its units, so the agents drive the browser the run chose.
Operational notes
Section titled “Operational notes”- Agents dial out: only the orchestrator’s gRPC and HTTPS ports need to be reachable.
- The orchestrator needs datastore connectivity to persist hierarchy-format runs; agents do not persist anything locally.
- Suite directories and functional results travel to and from the agents in 1 MiB chunks, so
their size is not limited by gRPC’s message size. For a large suite directory, raise
timeouts.readso the upload ofPOST /execute/runorPOST /execute/loadhas time to finish.