Outputs & Reports
What each runner emits and how to consume it in CI and locally.
v2 runner (kis test run)
Section titled “v2 runner (kis test run)”By default no files are written: results stream to the console and the exit code is 0
(all passed) / 1 (any failure). Four outputs are opt-in: JUnit XML, an HTML report, a results
database, and the files steps produce.
JUnit XML: --junit <path>
Section titled “JUnit XML: --junit <path>”Standard JUnit XML for CI ingestion (GitLab/GitHub test reports, Jenkins). Each test case’s
system-out lists its steps with their status, duration and summary, and what scripts logged;
a dry run marks its tests skipped:
kis test run -t tests/ --junit .kis/test/junit.xmlHTML report
Section titled “HTML report”--html <path> writes the run as one page to read in a browser:
kis test run -t tests/ --html .kis/test/report.html --artifacts .kis/test/artifactsThe page shows the run’s counts, then each suite and test. For each step it shows the summary, the error, every assertion with what was expected and what came back, what the step sent and received (HTTP request and response, CLI command and output, SQL statement and rows, gRPC payloads, script source, stream messages), its log, the variables it extracted, and the files it kept. Screenshots show inline, and a browser step’s errors are shown as Playwright wrote them.
Failed tests open by default and passing ones are collapsed; Failures only hides the
passing tests. before_all and after_all steps are listed with their suite, and a failing
before_each or after_each step too. Bodies and outputs longer than 64 KB are shortened on
the page; the --db store keeps them whole.
The page is self-contained and follows the system’s light or dark setting. Kept files are
linked relative to the page, so keep --artifacts beside it (as above) and move the two
together; a CI job that publishes both as one artifact gives a browsable report.
DuckDB: --db <path>
Section titled “DuckDB: --db <path>”Full structured results for querying and cross-run analysis:
kis test run -t tests/ --db .kis/test/results.dbTables:
| Table | Contents |
|---|---|
runs | one row per run: status, run_type, totals (total_tests, passed_tests, failed_tests, skipped_tests, errored_tests), duration_ms |
suite_runs | per-suite rollups, parent-linked (parent_suite_id) |
test_runs | per-test results |
step_runs | per-step results |
assertion_results | per-assertion: expression, passed, field, operator, expected, actual |
http_exchanges | wire evidence per HTTP step: dns_ms, connect_ms, tls_ms, ttfb_ms, transfer_ms, total_ms, request_size, response_size, redirects |
sql_executions | database step evidence: driver, statement, params, row_count, rows_affected, columns, duration_ms. The connection string is deliberately not stored: it carries credentials and a results file gets copied around |
grpc_exchanges | gRPC step evidence: address/service/method, request and response payloads, status |
script_executions | script step evidence including log_output (the console and JUnit also show it) |
stream_executions | one row per ws:/sse: step: protocol, url, connected, status_code, message_count (everything that arrived, even past the retention bound), duration_ms, error |
stream_triggers | one row per fired trigger: seq, kind (http/script), name (its save_as), method, url, status_code, message, duration_ms, error. A push at offset_ms N means nothing without the trigger it should be N after |
stream_messages | every retained message: seq, offset_ms, kind (text/binary, or the SSE event name), event_id, data. offset_ms is milliseconds since the connection opened, which is what makes “did the push arrive within 200ms of the trigger” answerable after the fact |
captured_variables | variables a step extracted or published, with the scope they reached (test, suite, run) |
step_evidence | evidence of a step type with no table of its own (such as playwright:), as JSON under its protocol name |
step_artifacts | files a step kept (artifacts:, Playwright reports, traces and screenshots): name, path of the kept copy, media type, size |
load_stats | load aggregates at three scopes: run, phase and scenario, each with counts, rps, error_rate, peak_vus and full latency/TTFB distributions |
route_coverage | one row per route the service registers: method, path (the template), source (file:line), hits, statuses, covered, and happy_path_only: exercised but never seen to fail, which no percentage shows |
code_coverage | Go statement coverage per file, from a service built with go build -cover |
coverage_summary | the headline numbers, one row per kind (routes, statements), so a trend is a query rather than a re-aggregation |
load_windows | the run as a trend: one row per sample_every interval with its own latency distribution, rps and error_rate. An end-of-run aggregate averages a degradation away; this is where it stays visible |
load_resources | the system under test’s own metrics on the same timeline: monitor, window_index, metric, value. Joins to load_windows on window_index, which is what lets one query say whether latency rose because the service was leaking, throttled, or simply given more work |
load_errors | the failure breakdown: scenario, label (step + assertion, with varying values stripped so like failures group), count. Individual iterations are not stored, so this is the only record of why a run had the error rate it had |
load_thresholds | each threshold’s metric, operator, expected, actual and outcome |
Load runs write to the same runs table as functional ones, with run_type = 'load', so
both list and compare together.
Individual load iterations are never stored: a one-second run at 4 VUs produces ~30,000, and the aggregates are what a comparison reads.
Query with the DuckDB CLI:
duckdb .kis/test/results.db "SELECT s.step_name, avg(x.total_ms) AS avg_ms, avg(x.ttfb_ms) AS avg_ttfb FROM step_runs s JOIN http_exchanges x ON x.step_run_id = s.id GROUP BY 1 ORDER BY 2 DESC"http_exchanges also stores full request/response headers (JSON) and bodies per step, so a
failing run can be forensically replayed from the DB alone.
Notes:
- DuckDB is single-writer: don’t point two concurrent runs at the same file.
redirectscounts the redirects followed to reach the final response.- The database is append-only: runs accumulate so trends stay queryable.
Comparing runs
Section titled “Comparing runs”Rather than remembering the schema, three commands wrap the questions worth asking:
kis test runs list --db .kis/test/results.db # what has run, newest firstkis test runs list --type load --last 5 # just load runskis test runs show <run-id> # one run in detailkis test runs compare <base-id> <new-id> # what changedkis test runs compare --last # the two most recentRun ids may be abbreviated to any unique prefix.
Load comparison reports the latency percentiles, throughput and error rate side by side, and
says whether each moved in a good or bad direction: rps going up is an improvement while
p95 going up is not, and the report knows the difference:
METRIC BASE NEW DELTAp95 latency ms 0.331 0.333 +0.002 (+0.6%) worserps 31493.888 31605.059 +111.171 (+0.4%) betterFunctional comparison names what changed rather than reporting a count, since “3 failed” does not tell you whether it is the same three:
regressed (1) ✗ beta (failed)
new (1) + gammaTests are matched across runs by their table group where they have one, so a data-driven test compares as a single entity even when its fixture changes size.
Console
Section titled “Console”Colored per-test/per-step streaming output; --no-color disables ANSI (auto-disabled when
stdout isn’t a TTY). Diagnostics (env discovery, db/junit paths) go to stderr, results to
stdout.
Each step prints one line: its status, name, duration and a summary of what it did (GET <url> → 200, <command> → exit 0, <driver> <statement> → 3 row(s), <service>/<method> → 0). A step that was retried shows its attempt count. Failed assertions and errors follow on
their own lines, then anything a script logged with test.log, then the kept files of a step
that failed. The run summary ends with the artifacts directory when any step kept files.
Artifacts: --artifacts <dir>
Section titled “Artifacts: --artifacts <dir>”Files steps produce are kept in a directory per run: <dir>/<run-id>/<test>-<id>/<NN>-<step>/,
with attempt-N/ added for a retried test. Without --artifacts the directory is under the
system temporary directory. A step keeps files it lists under artifacts: (see
Test definitions); a playwright: step keeps its report,
traces, screenshots and videos. JUnit lists each kept file as [[ATTACHMENT|<path>]] in the
test case’s system-out, the convention Jenkins and GitLab read. Load runs keep no artifacts.
Tests mode (kis test tests)
Section titled “Tests mode (kis test tests)”- Log file:
<prefix>-<RFC3339Nano timestamp>.login the working directory (--prefix, defaulttest), mirrored to stdout.-s/--logresponsesincludes response bodies;-n/--loglinenumbersadds line numbers. - Summary table after the run with canonical columns
TEST PLAN / TEST SCENARIO / TEST CASE / TEST STEP / STATUS / DURATION(✅ pass/❌ fail). Data-driven rows get-RC<n>suffixes when there are ≥ 2 records. - Load mode adds a status line every 100 ms and per-phase summary tables: see Load testing.
- Exit code
1on any failure.