Skip to content
Talk to our solutions team

Server Modes

For horizontal scale, run Testing as a control plane (orchestrator) plus workers (agents). The orchestrator spreads two kinds of run across its agents: a functional run of your suites (kis test run --orchestrator) and a load run (kis test load --orchestrator).

gRPC (agents dial out)
┌─────────┐ register + heartbeat + work ┌──────────────┐
│ agent │ ───────────────────────────────▶ │ orchestrator │◀── HTTPS API (JWT)
└─────────┘ ◀─────────────────────────────── └──────────────┘
runs its share dispatches, results returns results
of each run to the submitter
  • The orchestrator runs a gRPC server (agents register and heartbeat) and an HTTPS API for submitting runs and reading results.
  • Agents dial the orchestrator as gRPC clients (only the orchestrator needs an inbound gRPC port), run the work they are handed, and send results back over the same channel. Each agent also serves a minimal local HTTPS endpoint: GET /ready, GET /health.

Both commands read a .test.yaml config file (-f/--config) plus the shared infra flags (--port, --sid, --config_url, service-discovery and TLS cert flags; see CLI Reference).

Terminal window
test.svc orchestrator -f orchestrator.test.yaml

Key config:

KeyRequiredMeaning
grpcyesgRPC server port (startup fails without it)
host / portHTTPS API bind address
timeouts.read / timeouts.write / timeouts.readheaderserver timeouts (seconds or duration strings; defaults 15s/15s/5s)
orchestrator.queue.maxdepthwork-queue depth
serviceidservice identity
loadtenantstenants to warm-load at boot
grpccredsgRPC credentials
localagentoptional in-process agent: id, serverurl, grpc, heartbeatduration (default 15s), infrahost, grpcaddress, httpaddress

TLS comes from the container-registered service certificates (--svccert/--svckey). Shutdown is graceful: SIGINT/SIGTERM drains agent and orchestrator work with 30s timeouts.

All routes are JWT-secured.

RouteMethodPurpose
/execute/runPOSTstart a functional run of submitted suites across the registered agents
/execute/run/:idGETthe functional run’s status, and each shard’s results when it is done
/execute/loadPOSTstart a load run of a resolved profile across the registered agents
/execute/load/:idGETthe load run’s status, and the merged result when it is done
/list/testsGETteststeps/testcases/testscenarios/testplans known for the tenant (hierarchy format)
/execute/testsPOSTfan a filtered hierarchy-format test set out to agents
/list/test/runsGEThierarchy-format run history with metrics and logs
/agentsGETregistered agents
/workflow/queueGETqueue status
/workflow/queue/:idDELETEcancel queued work
/ready, /healthGETprobes

POST /execute/run and POST /execute/load answer 202 with the run’s id as soon as the run starts:

{"run_id": "6f1c…", "status": "running"}

GET /execute/run/:id (or /execute/load/:id) answers with the run. status is running until the run ends, then done with its result, or failed with an error saying why the run could not happen. A finished run stays available for an hour. The clients ask every two seconds, so a run lasts as long as it needs to while each request stays short.

Terminal window
kis test run -t tests/ --orchestrator https://orch.internal:8443 --token "$TALOS_TOKEN"

Your machine loads the suites exactly as a local run does, then:

  1. Splits the run into units. The root suite’s own tests are one unit, or one unit per test when the root suite is parallel: true. Each child suite of the root, with everything under it, is one unit. A sequential suite’s tests stay together, because a test may use what an earlier one published.
  2. Runs the root suite’s before_all here, once. What it extracted and published goes to the agents with the units.
  3. Packs the suite directory and submits it with the units. node_modules, vendor, .git, .artifacts, test-results and playwright-report are left out. When the suites read files outside --tests (a shared fixtures folder, an import from a sibling directory), pass --bundle-root to send the directory that contains both.
  4. The orchestrator spreads the units across its agents by test count, largest first onto the least loaded agent. --max-agents N caps how many are used.
  5. Each agent unpacks the directory and loads the suites from it, then runs its units with the same engine as a local run, starting from what before_all left. Every step type runs, playwright: included.
  6. The results come back: everything each unit reported, and the files its steps kept. Your machine passes them through its own reporters and store, so the console output, --junit, --html, --db and --artifacts match a local run of the same suites. Then it runs the root suite’s after_all, once.

A unit whose agent reports an error, or does not report within --timeout (default 30m), has each of its tests reported errored with the reason, so the run fails.

A value published with export: run reaches the tests of its own unit. Units may run on different agents, so publish what every unit needs from the root suite’s before_all.

POST /execute/run body, as kis test run --orchestrator sends it:

{
"run_id": "optional; one is minted if omitted",
"bundle": "<the suite directory, a base64 gzipped tar>",
"tests": "path of --tests within the bundle",
"units": [{"key": ".", "tests": 4}, {"key": "orders", "tests": 12}],
"setup": {"scope": {"token": "…"}, "run": {}, "suite": {}},
"tags": ["smoke"],
"exclude_tags": [],
"max_agents": 0,
"timeout_seconds": 1800
}

When done, result.shards lists each shard: shard, shards, host, the units it ran, its packed results, and an error when it could not run them.

POST /execute/load body:

{
"run_id": "optional; one is minted if omitted",
"profile": { "...": "a resolved load profile, tests already expanded" },
"bundle": "<the suite directory, a base64 gzipped tar>",
"bundle_root": "/home/you/project/tests",
"max_agents": 0,
"grace_seconds": 120
}

You rarely build this by hand: kis test load --orchestrator <url> resolves the suite and sends it with the suite directory, packed (bundle, and bundle_root, where the directory is on your machine). Each agent unpacks the directory and maps the profile’s paths onto its copy, so the files a step reads are there. Agents need no deployment step when your tests change, and a suite with a very large table: makes a larger dispatch.

Shards start together. Each agent prepares its shard and reports ready; the orchestrator waits for every dispatched shard, up to two minutes, then starts all that are ready at once. A shard that failed to prepare or was not ready in time is listed with the reason. Each shard also carries a deadline derived from the profile’s own duration, so generators stop when the run should have ended.

When done, result holds the merged result plus one entry per shard:

{
"result": { "Overall": {"...": "..."}, "Thresholds": ["..."] },
"shards": [
{"shard": 1, "shards": 3, "host": "gen-a", "contributed": true, "error": ""},
{"shard": 2, "shards": 3, "host": "gen-b", "contributed": false, "error": "not ready to start within 2m0s"}
]
}

Every shard is listed whether it contributed or not, and a run that lost a generator exits 1: less load reads as a faster service.

The machine that submits the run runs the suite’s before_all once, before dispatch, and its after_all once, after the merge; what setup extracted and published travels to the agents with the profile. Agents run neither, so a setup that seeds a fixed row or claims a unique name runs once.

GET /list/test/runs query params: id, filter, page (default 1), pagesize (default 10, max 100), metricfields, logfields. filter is a condition over the testrun fields, such as status = 'failed'; one that does not parse is answered 400.

POST /execute/tests body (hierarchy format):

{
"type": "testplan",
"pattern": "api-smoke",
"vars": {"baseurl": "https://api.staging.example.com"},
"logresponse": false,
"workercount": 2
}

Splits the filtered tests across workercount agents; responds {"message": "test runs started successfully", "status": {"runid": "..."}}.

Terminal window
test.svc agent -f agent.test.yaml

Key config:

KeyRequiredMeaning
serverurlyesorchestrator gRPC address to dial
sidagent instance id
grpcagent’s own gRPC port
heartbeatdurationdefault 15s
infrahost, grpcaddress, httpaddressaddresses advertised on registration
tenantstenant keys this agent serves
exclusivereserve the agent for its tenant list
max_concurrent_tasksconcurrency cap
host / port + timeouts.*the local /ready + /health HTTPS server
loadtenants, grpccredsas for the orchestrator

The agent takes three kinds of work:

  • Units of a functional run from /execute/run: it unpacks the suite directory into a temporary directory, loads the suites, runs its units, and sends back what they reported and the files their steps kept. The directory is removed afterwards.
  • A load shard from /execute/load: it unpacks the suite directory, prepares its share of the profile (users and rates divided across the agents, durations not), reports ready, starts when the orchestrator says so, and returns a mergeable snapshot; see Load testing.
  • A hierarchy-format run from /execute/tests (run id, environment, tests, tenant): it runs and streams results back. Metric records become testrunmetric rows, log lines become testrunlog rows, and a completion or error message closes out the scenariorun.

An agent runs playwright: steps with its own Node, Playwright and browser; the suite directory it receives has no node_modules. Install Node 18 or later and @playwright/test on the agent, and set NODE_PATH to the global node_modules directory so Node finds Playwright from any directory. For Obscura, the default browser, put obscura and obscura-worker on the agent’s PATH (or set TALOS_OBSCURA); for a browser Playwright launches, install it with npx playwright install:

Terminal window
npm install -g @playwright/test
export NODE_PATH="$(npm root -g)"
curl -L https://github.com/h4ckf0r0day/obscura/releases/download/v0.2.3/obscura-x86_64-linux.tar.gz | tar xz -C /usr/local/bin
npx playwright install --with-deps chromium # only for runs that choose chromium

The run’s --browser travels with its units, so the agents drive the browser the run chose.

  • Agents dial out: only the orchestrator’s gRPC and HTTPS ports need to be reachable.
  • The orchestrator needs datastore connectivity to persist hierarchy-format runs; agents do not persist anything locally.
  • Suite directories and functional results travel to and from the agents in 1 MiB chunks, so their size is not limited by gRPC’s message size. For a large suite directory, raise timeouts.read so the upload of POST /execute/run or POST /execute/load has time to finish.