Skip to content
Talk to our solutions team

Load Testing

Testing reuses your functional test definitions as the load workload: the suites kis test run runs are executed repeatedly by concurrent virtual users, phase by phase. You write a load profile describing phases and thresholds and point it at those suites.

SuitesCommand
v2 suiteskis test load -t <dir> --profiles <profile.yaml> --profile <name>
hierarchy-format suiteskis test tests -p <dir> --load <profile.yaml> --profile <name>

v2 load: the same suites you test functionally

Section titled “v2 load: the same suites you test functionally”

There is no separate workload format. The suites kis test run already runs become the load workload, so a suite you trust functionally is the one you put under load.

Terminal window
kis test load -t tests/ --profiles profile.yaml --profile stress
# profile.yaml: lives beside your tests; the suite loader skips it
profiles:
stress:
phases:
- type: ramp_up
duration: 30s
start_users: 1
end_users: 50
- type: constant
duration: 2m
users: 50
- type: ramp_down
duration: 30s
start_users: 50
end_users: 0
thresholds:
- metric: p95_latency_ms
operator: lt
value: 500
- metric: error_rate
operator: lt
value: 0.01
scenarios: # optional weighting; default is uniform
- test: checkout
weight: 3

A loadprofiles: file written for tests-mode also parses here, its warmup/sustained/cooldown becoming three constant phases: trying the v2 runner does not start with a rewrite.

A virtual user runs one test per iteration, then picks again. That is the difference from tests-mode, where a pass covers the entire collection.

It matters most for data tables. A test with table: becomes one scenario cycling its rows, one row per iteration, so a 500-row fixture gives a VU 500 different row values rather than multiplying your traffic 500-fold. Volume is set by VUs and duration alone.

Every suite’s before_all runs once, root suite first, before any VU starts. The variables it extracts and everything it publishes (export: run, test.set_run_var) are merged into the variables every iteration inherits. That is how you load test anything authenticated: log in once, hammer with the token.

after_all runs once at the end, deepest suite first, on its own context, so teardown still happens after an interrupted run.

Each iteration runs against the suite its test was declared in: that suite’s variables (rendered over its parents’, as in a functional run), and the same before_each / after_each chain a functional run would wrap it in. Their cost is part of the measured latency: they are steps of the journey, and excluding them would report a latency no user experiences.

Each virtual user has its own memory, starting from what setup published. What an iteration publishes (export: run, test.set_run_var) is visible to that virtual user’s later iterations, like one user’s session, and never to another’s.

Browser steps (playwright:) do not run under load: a browser per virtual user measures the load generator rather than the service. A profile that reaches one is refused before it starts. Load test the API the page calls, and exclude browser tests with --exclude-tags.

A failing before_all aborts the run rather than being counted as errors. Hammering a service whose precondition never happened measures nothing, and a 100% error rate is a broken harness, not a finding.

Two families. The difference is not cosmetic.

Closed-loop: a fixed population of users, each waiting for its response before sending again:

TypeFields
constantusers, duration
ramp_upstart_users, end_users, duration
ramp_downstart_users, end_users, duration

Open-loop: iterations scheduled on a clock regardless of whether earlier ones finished:

TypeFields
constant_raterate (iterations/sec), max_vus, duration
ramp_ratestart_rate, end_rate, max_vus, duration

Reach for open-loop when you need to hold a target RPS or model a traffic spike. A VU phase cannot do either: when the service slows, its users slow with it, so the offered load falls and the run quietly stops applying the pressure you asked for.

max_vus is required on a rate phase and bounds concurrency. Iterations that find no free worker are dropped and counted: dropping is the finding (the target could not absorb the rate), and queueing them instead would silently turn the test back into a closed loop:

iterations 4332 failed 86 peak VUs 18 rps 866.1
dropped 3287 — the target could not absorb the requested arrival rate

Ramps are linear and sampled every 100 ms. A ramp_down sheds VUs between iterations, never mid-request: an interrupted request would count as a failure the service did not cause. For the same reason an iteration interrupted by the phase deadline is discarded rather than measured: it describes your shutdown, not the service.

profiles:
realistic:
think_time: 2s
think_jitter: 0.3 # ±30%
phases: [...]

Pauses a VU between iterations. Without it a VU loops as fast as the service answers, which is a stress shape rather than a traffic shape. Jitter matters: without it every VU that started together stays in lockstep and arrives in waves a real population does not produce.

This is what makes a load run a gate. Every threshold is evaluated against the aggregated metrics; if any misses, the process exits 1.

Metrics available to a threshold: p50_latency_ms, p90_latency_ms, p95_latency_ms, p99_latency_ms, min_latency_ms, avg_latency_ms, max_latency_ms (every iteration); avg_latency_passed_ms, max_latency_passed_ms, p50_latency_passed_ms, p90_latency_passed_ms, p95_latency_passed_ms, p99_latency_passed_ms (passed iterations only); p95_ttfb_ms, p99_ttfb_ms; error_rate (0–1), rps, iterations, passed, failed, bytes_in, bytes_out, dropped, dropped_rate (0–1). Operators: lt/<, lte/<=, gt/>, gte/>=, eq/==. Add scenario: <test name> to scope a threshold to one scenario.

The names are exact: p95 is not one of them, p95_latency_ms is. A profile whose threshold names a metric or operator that does not exist fails to load, before the run starts, with the nearest valid name.

On a rate phase, gate on dropped_rate. Latency alone will not catch a saturated run: the iterations that never went out have no latency, so a target that absorbed a third of the requested rate can pass every percentile threshold. dropped_rate lt 0.01 reads as “we sustained at least 99% of the rate we asked for”. dropped and dropped_rate are run-scope only: the drop happens in the scheduler, before a scenario is picked, so there is no per-scenario answer.

iterations 31581 failed 0 peak VUs 4 rps 31479.7
error rate 0.00%
latency ms min 0.1 avg 0.1 p50 0.1 p90 0.3 p95 0.3 p99 0.5 max 2.2
cheap iters 15883 p95 0.3ms err 0.00%
compute iters 15698 p95 0.4ms err 0.00%
✓ pass error_rate lt 0.01 — actual 0.000
✓ pass p95_latency_ms lt 500 — actual 0.334

Latency is milliseconds here, unlike the tests-mode tables which print nanoseconds. Latency covers every iteration, failed ones included: an iteration that failed a duration_ms < 500 assertion is exactly the slow one a latency gate exists to catch. When some iterations failed, a second line gives the same percentiles over the passed ones only, which the *_latency_passed_ms thresholds read. TTFB covers iterations whose HTTP step reported one; a script or database scenario has no first byte.

An iteration’s latency is the whole test: its before_each, every step, its after_each. rps counts iterations per second over the time the phases ran; setup and teardown are not part of it. An iteration still in flight when its phase ends finishes and is counted in that phase, so a phase overruns by at most one iteration.

A run also prints a failure breakdown:

failures
148 100.0% place-order post-order: status_code == 200

Like failures collapse to one line: the label carries the step and the assertion, never the actual values, so ten thousand occurrences of one problem do not become ten thousand rows. Individual iterations are deliberately not stored (a one-second run at 8 VUs produces tens of thousands), so this aggregate is the only record of what went wrong. It is persisted to load_errors.

While the run is in flight it reports itself every second to stderr:

0:42 phase 2/3 constant vus 8/8 iters 12480 fail 31 rps 297.1

Automatically off when stderr is not a terminal, so CI logs stay readable; --no-progress turns it off explicitly.

Latency is recorded in an HDR-style histogram rather than by keeping every sample, because keeping them costs about 537MB per million iterations and a two-hour soak simply runs out of memory. Percentiles are therefore accurate to roughly 1.6% worst case, about half that typically: far below the run-to-run variance of any real service. min, max and the mean are tracked exactly.

The histogram is also what makes the next section possible.

profiles:
soak:
sample_every: 10s
phases: [{type: constant, duration: 6h, users: 50}]

Each interval keeps its own latency distribution, not a running total, and the summary shows the shape:

trend
p95 ▁▁▂▂▃▄▅▆▇█ 18.2 → 184.6 ms (+914%) over 36 windows

This is the only thing that answers a soak’s question. An end-of-run p95 averages the first healthy hour together with the last degraded one and reports something unremarkable: in testing, a run whose latency went from 6.6ms to 62ms reported an aggregate of 62 and said nothing about the 6.6.

Stored in load_windows, one row per interval, so the trend survives the run and runs compare can read it.

Resource metrics: what the service itself was doing

Section titled “Resource metrics: what the service itself was doing”

A load run measures the wire. It cannot see that the service was leaking, throttled, or swapping, and those explain most of what a load test finds. Run the monitor on the target:

Terminal window
# on the machine under test
kis test monitor --port 8080 --collect process,host
# in the profile
profiles:
soak:
sample_every: 10s
monitors:
- name: api
url: "http://api-host:9099"
collect: [process.rss_bytes, process.cpu_pct, host.cpu_pct]
resources
api process.rss_bytes 14.1 MB → 3.25 GB peak 3.25 GB ↑ +23468% over the run
api process.cpu_pct 38.7% → 40.9% peak 42.7%

That run’s latency was flat at 4.5ms with zero errors, every threshold passed, while the service consumed 3.25GB in ten seconds. No amount of wire-level measurement would have found it.

Resource samples carry the same window index as the latency series, so one query answers “did latency rise because it was leaking, throttled, or just given more work”:

SELECT w.offset_s, w.rps, w.p95_latency_ms,
max(CASE WHEN r.metric='process.rss_bytes' THEN r.value END)/1048576 AS rss_mb
FROM load_windows w
LEFT JOIN load_resources r
ON r.run_id = w.run_id AND r.window_index = w.window_index
GROUP BY 1,2,3 ORDER BY 1;

kis test monitor --list prints the catalogue. Four groups, selectable whole (--collect process,host) or by individual name:

GroupMetrics
processrss_bytes, vsz_bytes, cpu_pct, threads, open_fds
hostcpu_pct, load1, mem_used_bytes, mem_total_bytes, disk_used_pct, net_rx_bytes, net_tx_bytes
cgroupmem_limit_bytes, mem_usage_bytes, cpu_quota, cpu_throttled_pct
goheap_alloc_bytes, heap_sys_bytes, goroutines, gc_count, gc_pause_ms: needs --endpoint pointing at /debug/vars

process.rss_bytes is the one that matters most, because this platform runs v8go and duckdb through cgo and neither allocates on the Go heap. A leak can be invisible to go.heap_* and obvious in RSS. Collect both: the pair says whether growth is Go-side or cgo-side, which decides where to look.

cgroup.cpu_throttled_pct explains a service that is slow without being busy. A throttled container burns its quota and waits; on process CPU alone it looks idle while every request queues.

--pid, --name (newest match) or --port (whatever is listening). Name and port are re-resolved on every reading, so a service that restarts mid-soak keeps being measured instead of reporting a dead pid’s last values.

The run owns the clock, so a sample it pulled belongs unambiguously to the window it pulled for. A pushing monitor arrives on its own schedule and ends up half a window out of step with the latency it exists to explain. It also means the monitor holds no buffer: a generator that dies leaves nothing behind.

A caller may narrow what it asks for (collect: in the profile) and can never widen it. The monitor’s own flags are the authority on what a host exposes.

An unreachable monitor costs one series, not the run, and warns once rather than once per interval.

The same metrics through an OpenTelemetry Collector

Section titled “The same metrics through an OpenTelemetry Collector”

If you already run an OpenTelemetry Collector, it can report the same names through its hostmetrics receiver:

Terminal window
kis test monitor --collect process,host --name orders-service --otel

prints a receiver and filter configuration for exactly that selection, filtered to the same metric names, so a run measured either way is comparable. It also names the metrics that come from elsewhere: cgroup metrics from the monitor itself, and go.* from the service’s own OpenTelemetry SDK.

Horizontal scale: several generators, one report

Section titled “Horizontal scale: several generators, one report”

One machine saturates before most targets do. Run the same profile on several, then merge:

Terminal window
box1$ kis test load -t tests/ --profiles load.yaml --profile stress --shard 1/3 --out box1.json
box2$ kis test load -t tests/ --profiles load.yaml --profile stress --shard 2/3 --out box2.json
box3$ kis test load -t tests/ --profiles load.yaml --profile stress --shard 3/3 --out box3.json
any$ kis test load merge box*.json --db .kis/test/results.db

--shard N/M divides users and rates across generators. It does not divide durations: three machines running a two-minute profile produce two minutes of load at three times the rate. Remainders go to the low-numbered shards, so 10 users across 3 generators is 4/3/3 rather than a quietly missing user.

You cannot combine generators’ summaries, the average of two p95s is not a p95, so each writes the distribution it measured and merge does the reading. Counts add, distributions add bucket by bucket (exact), peak concurrency adds (the generators run at the same time), and rates divide by the longest generator’s own load time, so the generators’ clocks are never compared. The trend (sample_every) and resource samples merge too, window by window, and the failure breakdown ranks every distinct failure across generators.

Thresholds are not divided. They are statements about the service under the whole load, so each shard carries them unchanged and they are judged after the merge: a per-shard verdict would gate on a third of the traffic. This is why a sharded run should be gated on the merge, not on the individual shards.

A merge that is missing a shard says so, and exits 1: it covers part of the requested load, and smaller numbers would read as a faster service.

merged 2 generator(s): box1 1/3, box3 3/3; WARNING: shard(s) 2 missing, this is 2/3 of the intended load

With --shard, each generator runs the suite’s before_all and after_all itself, so a setup that seeds a fixed row or claims a unique name needs to be idempotent.

Orchestrated: let the orchestrator do the sharding

Section titled “Orchestrated: let the orchestrator do the sharding”

If you already run test.svc orchestrator and one or more test.svc agent processes, point the run at the orchestrator instead and it shards across whatever agents are registered:

Terminal window
kis test load -t tests/ --profiles load.yaml --profile stress \
--orchestrator https://orch.internal:8443 --token "$TALOS_TOKEN"
submitting to https://orch.internal:8443: the orchestrator shards across its registered agents
run 6f1c… submitted; waiting for it to finish
run 6f1c…: 3 generator(s): gen-a 1/3, gen-b 2/3, gen-c 3/3 (2m31s)

--max-agents N caps how many are used; the default is all of them.

Your machine resolves the suite, and the agents run it from a copy. The expanded tests travel with the request, and so does the suite directory, packed: an agent unpacks it and runs every step against its copy, so a step that reads a fixture, runs a script the suite ships, uploads a file or presents a certificate finds it on the agent. Agents need no deployment step when your tests change, and a run cannot depend on several machines agreeing about a CSV. When the suites read files outside --tests, pass --bundle-root to send the directory that holds both; node_modules, vendor, .git and earlier run outputs are left out. A suite with a very large table: makes a large dispatch: the rows are sent, once per agent.

Setup and teardown run once, on your machine. before_all runs where you submitted the run, before dispatch, and what it extracted and published travels to the agents with the profile; after_all runs there once every agent has finished. Agents run neither.

Every shard’s outcome is listed whether it contributed or not, and a run that lost a generator exits 1: it produced smaller numbers, which read as a service that got faster.

run 6f1c…: 2 generator(s): gen-a 1/3, gen-c 3/3; WARNING: 1 did not contribute: gen-b 2/3 (not ready to start within 2m0s)

The gate is judged on the merged numbers, for the same reason as the manual path: each agent only ever saw its own share.

Shards start together. Each agent prepares its shard and reports ready, and the orchestrator starts every shard that is ready at once, so the merged run describes one load applied from one moment. A shard that failed to prepare, or was not ready within two minutes, is listed with the reason.

The run is submitted, then polled: the orchestrator answers the submission with the run’s id, and the client asks for the run every two seconds until it is done. Each shard also carries a deadline derived from the profile’s own duration, so generators stop when the run should have ended.

--shard + merge--orchestrator
Needsnothing but the binary on each boxa running orchestrator + agents
Who starts the generatorsyou (for-loop, CI matrix, Ansible)the orchestrator
Mergingyou run load mergeautomatic
Good forad-hoc scale, CI runners you already havea standing load fleet

Suites in the plan/scenario/case hierarchy format run under load with kis test tests --load:

Terminal window
kis test tests -p <dir> --load profile.yaml --profile default

Their profiles declare warmup, sustained and cooldown phases, each with a duration and targetusers, and the run prints a summary table per phase. Thresholds and trends are part of kis test load, which runs v2 suites.

  • Closed-loop phases bound concurrency, not throughput. Throughput is roughly users divided by the iteration’s mean latency. To hold a request rate, use constant_rate or ramp_rate.
  • Give a rate phase enough workers. A rate phase needs about rate × the iteration’s p95 latency in seconds workers busy at once; set max_vus comfortably above that, and gate on dropped_rate so a run that could not keep up fails instead of reporting good latency for the requests it managed to send.
  • Add think time for traffic, leave it out for stress. think_time with some jitter models users reading a page; without it, virtual users send as fast as the service answers.
  • Watch the generator. When the machine generating load runs out of CPU, its measurements describe itself. Spread the run with --shard or --orchestrator before that point.
  • Gate on the merged numbers of a sharded run, never on one shard’s.