Load Testing
Testing reuses your functional test definitions as the load workload: the suites
kis test run runs are executed repeatedly by concurrent virtual users, phase by phase. You
write a load profile describing phases and thresholds and point it at those suites.
| Suites | Command |
|---|---|
| v2 suites | kis test load -t <dir> --profiles <profile.yaml> --profile <name> |
| hierarchy-format suites | kis test tests -p <dir> --load <profile.yaml> --profile <name> |
v2 load: the same suites you test functionally
Section titled “v2 load: the same suites you test functionally”There is no separate workload format. The suites kis test run already runs become the load
workload, so a suite you trust functionally is the one you put under load.
kis test load -t tests/ --profiles profile.yaml --profile stress# profile.yaml: lives beside your tests; the suite loader skips itprofiles: stress: phases: - type: ramp_up duration: 30s start_users: 1 end_users: 50 - type: constant duration: 2m users: 50 - type: ramp_down duration: 30s start_users: 50 end_users: 0 thresholds: - metric: p95_latency_ms operator: lt value: 500 - metric: error_rate operator: lt value: 0.01 scenarios: # optional weighting; default is uniform - test: checkout weight: 3A loadprofiles: file written for tests-mode also parses here, its warmup/sustained/cooldown
becoming three constant phases: trying the v2 runner does not start with a rewrite.
What an iteration is
Section titled “What an iteration is”A virtual user runs one test per iteration, then picks again. That is the difference from tests-mode, where a pass covers the entire collection.
It matters most for data tables. A test with table: becomes one scenario cycling its rows,
one row per iteration, so a 500-row fixture gives a VU 500 different row values rather than
multiplying your traffic 500-fold. Volume is set by VUs and duration alone.
Setup and teardown
Section titled “Setup and teardown”Every suite’s before_all runs once, root suite first, before any VU starts. The
variables it extracts and everything it publishes (export: run, test.set_run_var) are
merged into the variables every iteration inherits. That is how you load test anything
authenticated: log in once, hammer with the token.
after_all runs once at the end, deepest suite first, on its own context, so teardown still
happens after an interrupted run.
Each iteration runs against the suite its test was declared in: that suite’s variables
(rendered over its parents’, as in a functional run), and the same before_each / after_each
chain a functional run would wrap it in. Their cost is part of the measured latency: they are
steps of the journey, and excluding them would report a latency no user experiences.
Each virtual user has its own memory, starting from what setup published. What an iteration
publishes (export: run, test.set_run_var) is visible to that virtual user’s later
iterations, like one user’s session, and never to another’s.
Browser steps (playwright:) do not run under load: a browser per virtual user measures the
load generator rather than the service. A profile that reaches one is refused before it
starts. Load test the API the page calls, and exclude browser tests with --exclude-tags.
A failing before_all aborts the run rather than being counted as errors. Hammering a
service whose precondition never happened measures nothing, and a 100% error rate is a broken
harness, not a finding.
Phases
Section titled “Phases”Two families. The difference is not cosmetic.
Closed-loop: a fixed population of users, each waiting for its response before sending again:
| Type | Fields |
|---|---|
constant | users, duration |
ramp_up | start_users, end_users, duration |
ramp_down | start_users, end_users, duration |
Open-loop: iterations scheduled on a clock regardless of whether earlier ones finished:
| Type | Fields |
|---|---|
constant_rate | rate (iterations/sec), max_vus, duration |
ramp_rate | start_rate, end_rate, max_vus, duration |
Reach for open-loop when you need to hold a target RPS or model a traffic spike. A VU phase cannot do either: when the service slows, its users slow with it, so the offered load falls and the run quietly stops applying the pressure you asked for.
max_vus is required on a rate phase and bounds concurrency. Iterations that find no free
worker are dropped and counted: dropping is the finding (the target could not absorb the
rate), and queueing them instead would silently turn the test back into a closed loop:
iterations 4332 failed 86 peak VUs 18 rps 866.1 dropped 3287 — the target could not absorb the requested arrival rateRamps are linear and sampled every 100 ms. A ramp_down sheds VUs between iterations,
never mid-request: an interrupted request would count as a failure the service did not cause.
For the same reason an iteration interrupted by the phase deadline is discarded rather than
measured: it describes your shutdown, not the service.
Think time
Section titled “Think time”profiles: realistic: think_time: 2s think_jitter: 0.3 # ±30% phases: [...]Pauses a VU between iterations. Without it a VU loops as fast as the service answers, which is a stress shape rather than a traffic shape. Jitter matters: without it every VU that started together stays in lockstep and arrives in waves a real population does not produce.
Thresholds and the exit code
Section titled “Thresholds and the exit code”This is what makes a load run a gate. Every threshold is evaluated against the aggregated metrics; if any misses, the process exits 1.
Metrics available to a threshold: p50_latency_ms, p90_latency_ms, p95_latency_ms,
p99_latency_ms, min_latency_ms, avg_latency_ms, max_latency_ms (every iteration);
avg_latency_passed_ms, max_latency_passed_ms, p50_latency_passed_ms,
p90_latency_passed_ms, p95_latency_passed_ms, p99_latency_passed_ms (passed iterations
only); p95_ttfb_ms, p99_ttfb_ms; error_rate (0–1), rps, iterations, passed,
failed, bytes_in, bytes_out, dropped, dropped_rate (0–1). Operators: lt/<,
lte/<=, gt/>, gte/>=, eq/==. Add scenario: <test name> to scope a threshold
to one scenario.
The names are exact: p95 is not one of them, p95_latency_ms is. A profile whose threshold
names a metric or operator that does not exist fails to load, before the run starts, with the
nearest valid name.
On a rate phase, gate on dropped_rate. Latency alone will not catch a saturated run:
the iterations that never went out have no latency, so a target that absorbed a third of the
requested rate can pass every percentile threshold. dropped_rate lt 0.01 reads as “we
sustained at least 99% of the rate we asked for”. dropped and dropped_rate are run-scope
only: the drop happens in the scheduler, before a scenario is picked, so there is no
per-scenario answer.
Reading the output
Section titled “Reading the output” iterations 31581 failed 0 peak VUs 4 rps 31479.7 error rate 0.00%
latency ms min 0.1 avg 0.1 p50 0.1 p90 0.3 p95 0.3 p99 0.5 max 2.2
cheap iters 15883 p95 0.3ms err 0.00% compute iters 15698 p95 0.4ms err 0.00%
✓ pass error_rate lt 0.01 — actual 0.000 ✓ pass p95_latency_ms lt 500 — actual 0.334Latency is milliseconds here, unlike the tests-mode tables which print nanoseconds.
Latency covers every iteration, failed ones included: an iteration that failed a
duration_ms < 500 assertion is exactly the slow one a latency gate exists to catch. When some
iterations failed, a second line gives the same percentiles over the passed ones only, which
the *_latency_passed_ms thresholds read. TTFB covers iterations whose HTTP step reported one;
a script or database scenario has no first byte.
An iteration’s latency is the whole test: its before_each, every step, its after_each.
rps counts iterations per second over the time the phases ran; setup and teardown are not
part of it. An iteration still in flight when its phase ends finishes and is counted in that
phase, so a phase overruns by at most one iteration.
A run also prints a failure breakdown:
failures 148 100.0% place-order post-order: status_code == 200Like failures collapse to one line: the label carries the step and the assertion, never the
actual values, so ten thousand occurrences of one problem do not become ten thousand rows.
Individual iterations are deliberately not stored (a one-second run at 8 VUs produces tens
of thousands), so this aggregate is the only record of what went wrong. It is persisted to
load_errors.
While the run is in flight it reports itself every second to stderr:
0:42 phase 2/3 constant vus 8/8 iters 12480 fail 31 rps 297.1Automatically off when stderr is not a terminal, so CI logs stay readable; --no-progress
turns it off explicitly.
Percentiles are approximate
Section titled “Percentiles are approximate”Latency is recorded in an HDR-style histogram rather than by keeping every sample, because
keeping them costs about 537MB per million iterations and a two-hour soak simply runs out of
memory. Percentiles are therefore accurate to roughly 1.6% worst case, about half that
typically: far below the run-to-run variance of any real service. min, max and the mean
are tracked exactly.
The histogram is also what makes the next section possible.
Trends: reading a run as a shape
Section titled “Trends: reading a run as a shape”profiles: soak: sample_every: 10s phases: [{type: constant, duration: 6h, users: 50}]Each interval keeps its own latency distribution, not a running total, and the summary shows the shape:
trend p95 ▁▁▂▂▃▄▅▆▇█ 18.2 → 184.6 ms (+914%) over 36 windowsThis is the only thing that answers a soak’s question. An end-of-run p95 averages the first healthy hour together with the last degraded one and reports something unremarkable: in testing, a run whose latency went from 6.6ms to 62ms reported an aggregate of 62 and said nothing about the 6.6.
Stored in load_windows, one row per interval, so the trend survives
the run and runs compare can read it.
Resource metrics: what the service itself was doing
Section titled “Resource metrics: what the service itself was doing”A load run measures the wire. It cannot see that the service was leaking, throttled, or swapping, and those explain most of what a load test finds. Run the monitor on the target:
# on the machine under testkis test monitor --port 8080 --collect process,host# in the profileprofiles: soak: sample_every: 10s monitors: - name: api url: "http://api-host:9099" collect: [process.rss_bytes, process.cpu_pct, host.cpu_pct] resources api process.rss_bytes 14.1 MB → 3.25 GB peak 3.25 GB ↑ +23468% over the run api process.cpu_pct 38.7% → 40.9% peak 42.7%That run’s latency was flat at 4.5ms with zero errors, every threshold passed, while the service consumed 3.25GB in ten seconds. No amount of wire-level measurement would have found it.
Resource samples carry the same window index as the latency series, so one query answers “did latency rise because it was leaking, throttled, or just given more work”:
SELECT w.offset_s, w.rps, w.p95_latency_ms, max(CASE WHEN r.metric='process.rss_bytes' THEN r.value END)/1048576 AS rss_mb FROM load_windows w LEFT JOIN load_resources r ON r.run_id = w.run_id AND r.window_index = w.window_index GROUP BY 1,2,3 ORDER BY 1;What it collects
Section titled “What it collects”kis test monitor --list prints the catalogue. Four groups, selectable
whole (--collect process,host) or by individual name:
| Group | Metrics |
|---|---|
process | rss_bytes, vsz_bytes, cpu_pct, threads, open_fds |
host | cpu_pct, load1, mem_used_bytes, mem_total_bytes, disk_used_pct, net_rx_bytes, net_tx_bytes |
cgroup | mem_limit_bytes, mem_usage_bytes, cpu_quota, cpu_throttled_pct |
go | heap_alloc_bytes, heap_sys_bytes, goroutines, gc_count, gc_pause_ms: needs --endpoint pointing at /debug/vars |
process.rss_bytes is the one that matters most, because this
platform runs v8go and duckdb through cgo and neither allocates on the
Go heap. A leak can be invisible to go.heap_* and obvious in RSS.
Collect both: the pair says whether growth is Go-side or cgo-side, which
decides where to look.
cgroup.cpu_throttled_pct explains a service that is slow without
being busy. A throttled container burns its quota and waits; on
process CPU alone it looks idle while every request queues.
Pointing it at a process
Section titled “Pointing it at a process”--pid, --name (newest match) or --port (whatever is listening).
Name and port are re-resolved on every reading, so a service that
restarts mid-soak keeps being measured instead of reporting a dead pid’s
last values.
Why it listens rather than pushes
Section titled “Why it listens rather than pushes”The run owns the clock, so a sample it pulled belongs unambiguously to the window it pulled for. A pushing monitor arrives on its own schedule and ends up half a window out of step with the latency it exists to explain. It also means the monitor holds no buffer: a generator that dies leaves nothing behind.
A caller may narrow what it asks for (collect: in the profile) and
can never widen it. The monitor’s own flags are the authority on what a
host exposes.
An unreachable monitor costs one series, not the run, and warns once rather than once per interval.
The same metrics through an OpenTelemetry Collector
Section titled “The same metrics through an OpenTelemetry Collector”If you already run an OpenTelemetry Collector, it can report the same
names through its hostmetrics receiver:
kis test monitor --collect process,host --name orders-service --otelprints a receiver and filter configuration for exactly that selection,
filtered to the same metric names, so a run measured either way is
comparable. It also names the metrics that come from elsewhere: cgroup
metrics from the monitor itself, and go.* from the service’s own
OpenTelemetry SDK.
Horizontal scale: several generators, one report
Section titled “Horizontal scale: several generators, one report”One machine saturates before most targets do. Run the same profile on several, then merge:
box1$ kis test load -t tests/ --profiles load.yaml --profile stress --shard 1/3 --out box1.jsonbox2$ kis test load -t tests/ --profiles load.yaml --profile stress --shard 2/3 --out box2.jsonbox3$ kis test load -t tests/ --profiles load.yaml --profile stress --shard 3/3 --out box3.json
any$ kis test load merge box*.json --db .kis/test/results.db--shard N/M divides users and rates across generators. It does not divide durations:
three machines running a two-minute profile produce two minutes of load at three times the
rate. Remainders go to the low-numbered shards, so 10 users across 3 generators is 4/3/3 rather
than a quietly missing user.
You cannot combine generators’ summaries, the average of two p95s is not a p95, so each
writes the distribution it measured and merge does the reading. Counts add, distributions add
bucket by bucket (exact), peak concurrency adds (the generators run at the same time), and
rates divide by the longest generator’s own load time, so the generators’ clocks are never
compared. The trend (sample_every) and resource samples merge too, window by window, and the
failure breakdown ranks every distinct failure across generators.
Thresholds are not divided. They are statements about the service under the whole load, so each shard carries them unchanged and they are judged after the merge: a per-shard verdict would gate on a third of the traffic. This is why a sharded run should be gated on the merge, not on the individual shards.
A merge that is missing a shard says so, and exits 1: it covers part of the requested load, and smaller numbers would read as a faster service.
merged 2 generator(s): box1 1/3, box3 3/3; WARNING: shard(s) 2 missing, this is 2/3 of the intended loadWith --shard, each generator runs the suite’s before_all and after_all itself, so a setup
that seeds a fixed row or claims a unique name needs to be idempotent.
Orchestrated: let the orchestrator do the sharding
Section titled “Orchestrated: let the orchestrator do the sharding”If you already run test.svc orchestrator and one or more test.svc agent processes, point
the run at the orchestrator instead and it shards across whatever agents are registered:
kis test load -t tests/ --profiles load.yaml --profile stress \ --orchestrator https://orch.internal:8443 --token "$TALOS_TOKEN"submitting to https://orch.internal:8443: the orchestrator shards across its registered agentsrun 6f1c… submitted; waiting for it to finishrun 6f1c…: 3 generator(s): gen-a 1/3, gen-b 2/3, gen-c 3/3 (2m31s)--max-agents N caps how many are used; the default is all of them.
Your machine resolves the suite, and the agents run it from a copy. The expanded tests
travel with the request, and so does the suite directory, packed: an agent unpacks it and runs
every step against its copy, so a step that reads a fixture, runs a script the suite ships,
uploads a file or presents a certificate finds it on the agent. Agents need no deployment step
when your tests change, and a run cannot depend on several machines agreeing about a CSV. When
the suites read files outside --tests, pass --bundle-root to send the directory that holds
both; node_modules, vendor, .git and earlier run outputs are left out. A suite with a very
large table: makes a large dispatch: the rows are sent, once per agent.
Setup and teardown run once, on your machine. before_all runs where you submitted the
run, before dispatch, and what it extracted and published travels to the agents with the
profile; after_all runs there once every agent has finished. Agents run neither.
Every shard’s outcome is listed whether it contributed or not, and a run that lost a generator exits 1: it produced smaller numbers, which read as a service that got faster.
run 6f1c…: 2 generator(s): gen-a 1/3, gen-c 3/3; WARNING: 1 did not contribute: gen-b 2/3 (not ready to start within 2m0s)The gate is judged on the merged numbers, for the same reason as the manual path: each agent only ever saw its own share.
Shards start together. Each agent prepares its shard and reports ready, and the orchestrator starts every shard that is ready at once, so the merged run describes one load applied from one moment. A shard that failed to prepare, or was not ready within two minutes, is listed with the reason.
The run is submitted, then polled: the orchestrator answers the submission with the run’s id, and the client asks for the run every two seconds until it is done. Each shard also carries a deadline derived from the profile’s own duration, so generators stop when the run should have ended.
Which one to use
Section titled “Which one to use”--shard + merge | --orchestrator | |
|---|---|---|
| Needs | nothing but the binary on each box | a running orchestrator + agents |
| Who starts the generators | you (for-loop, CI matrix, Ansible) | the orchestrator |
| Merging | you run load merge | automatic |
| Good for | ad-hoc scale, CI runners you already have | a standing load fleet |
Hierarchy-format suites under load
Section titled “Hierarchy-format suites under load”Suites in the plan/scenario/case hierarchy format run under load with
kis test tests --load:
kis test tests -p <dir> --load profile.yaml --profile defaultTheir profiles declare warmup, sustained and cooldown phases, each with a duration and
targetusers, and the run prints a summary table per phase. Thresholds and trends are part of
kis test load, which runs v2 suites.
Sizing guidance
Section titled “Sizing guidance”- Closed-loop phases bound concurrency, not throughput. Throughput is roughly users divided
by the iteration’s mean latency. To hold a request rate, use
constant_rateorramp_rate. - Give a rate phase enough workers. A rate phase needs about
rate× the iteration’s p95 latency in seconds workers busy at once; setmax_vuscomfortably above that, and gate ondropped_rateso a run that could not keep up fails instead of reporting good latency for the requests it managed to send. - Add think time for traffic, leave it out for stress.
think_timewith some jitter models users reading a page; without it, virtual users send as fast as the service answers. - Watch the generator. When the machine generating load runs out of CPU, its measurements
describe itself. Spread the run with
--shardor--orchestratorbefore that point. - Gate on the merged numbers of a sharded run, never on one shard’s.