Test Definitions
This is the schema for the v2 runner (kis test run). Tests are YAML, organized in
directories, with three levels:
Suite (directory or file) └── Test (unit of pass/fail) └── Step (one protocol action + assertions)Suites in the older steps:/testcases:/scenarios:/testplans: format run with
kis test tests; see the CLI Reference.
Discovery: directory = suite
Section titled “Discovery: directory = suite”kis test run -t <path> accepts a single YAML file (one flat suite) or a directory:
- Every directory is a suite; subdirectories become child suites, recursively.
- A file named
suite.yamlin a directory supplies suite-level config (name, variables, tags, data, databases, hooks, parallel). All other*.yaml/*.ymlfiles contribute only theirtests:. A test file that declaresvariables:,data:,databases:, or a hook at file level fails the load, with an error naming the file, the key, and the directory’ssuite.yaml. See Where suite-level keys live. - Files and subdirectories are processed in sorted order.
- Skipped: hidden entries (leading
.), directories starting with_(the convention for shared YAML pulled in viaimport:),node_modulesandvendor(package installs, such as a browser suite’s Playwright), and directories containing no YAML anywhere. - Multiple YAML files in one directory contribute their tests to the same suite.
tests/├── suite.yaml # root suite config: variables, tags, hooks├── env.yaml # env file (auto-discovered: not a test file)├── users/│ ├── suite.yaml # child suite; inherits root variables/tags│ ├── crud.yaml # tests│ └── permissions.yaml # tests└── orders/ └── checkout.yaml # child suite without its own suite.yamlComposition (import:)
Section titled “Composition (import:)”Files that declare a top-level import: block are expanded through yamlplus
(merge: / replace: / insert: / splice: verbs with selectors), with the file’s directory
as the base for relative references. Files without import: are parsed as plain YAML. Put
shared step blocks in a _-prefixed directory so they aren’t loaded as suites themselves. See
Reusing steps for the common case.
Suite schema
Section titled “Suite schema”Top-level keys of suite.yaml (or of a single-file suite):
| Key | Type | Default | Meaning |
|---|---|---|---|
name | string | directory/file basename | Suite identifier. |
tags | []string | — | Accumulate down the tree: effective tags = ancestor suite tags ∪ test tags. |
variables | map | — | Visible to every test in this suite and children; deeper suite overrides shallower. |
data | map[name]DataSource | — | Named data sources for table: iteration: see Data-driven tests. |
skip | bool or string | — | Skips every test in this suite and its child suites, with the reason; their hooks do not run. |
http | HTTPDefaults | — | Defaults for every HTTP request in this suite and its children: base_url, headers, timeout, follow_redirects, tls. See Suite defaults. |
parallel | bool | false | Run this suite’s tests concurrently. |
max_parallel | int | 0 = unbounded | Concurrency cap when parallel: true. |
tests | []Test | — | Tests declared in this file. |
before_all / before_each / after_each / after_all | []Step | — | Hooks; same shape as test steps. See below. |
Where suite-level keys live
Section titled “Where suite-level keys live”variables, data, databases, http, skip, and the four hooks are suite-scoped. Where they are read
depends on how the run is started:
| Entry point | Where suite-scoped keys are read |
|---|---|
-t dir/ | Only from each directory’s suite.yaml. |
-t file.yaml | From the file itself, which is the whole suite. |
In a directory run, a test file (any YAML file other than suite.yaml) that declares one of
these keys at file level fails the load:
load tests/products.yaml: suite-level keys in a test file are not applied when loading a directory; only suite.yaml is read for them: before_all: move it to tests/suite.yaml, or into the test's own steps:Move the key to the directory’s suite.yaml. A value that only one test needs can go on the
test instead: variables: on the test, or a setup step at the start of its steps:. A file
that has to run the same way under both entry points keeps its suite-scoped keys in
suite.yaml and only tests: in the file.
Test schema
Section titled “Test schema”| Key | Type | Required | Default | Meaning |
|---|---|---|---|---|
name | string | yes | — | Unique within the suite. |
tags | []string | no | — | Combined with suite-chain tags for --tags filtering. |
skip | bool or string | no | — | true, or the reason as a string. The test reports as skipped, with the reason, and nothing in it runs. |
only | bool | no | false | When any test in the run sets it, only those tests run. See Tags and filtering. |
variables | map | no | — | Merged over the suite snapshot before hooks/steps run. |
table | string | no | — | Data source name; runs the steps once per row. |
steps | []Step | yes | — | Executed in order. |
continue_on_error | bool | no | false | Keep running steps after a failure (reports all failures); also lets table iteration continue past a failing row. |
timeout | duration | no | none | Caps each attempt of the test: its before_each chain and its steps. A test that outlasts it errors with test timed out after <timeout>; after_each still runs. |
retries | int | no | 0 | Runs a test that failed or errored again from the start, with a fresh scope and cookie jar, up to this many more times. The result is the last attempt’s; the console prints ↻ <test> passed on attempt N so a flaky test stays visible, and --db stores the count in test_runs.retry_count. |
Step schema
Section titled “Step schema”Each step declares exactly one protocol key.
| Key | Type | Required | Default | Meaning |
|---|---|---|---|---|
name | string | yes | — | Step identifier. |
http / grpc / db / script / cli / ws / sse / playwright | object | one, or steps | — | Protocol config (below). sql: is accepted as the old name for db:. |
steps | []Step | instead of a protocol | — | Makes this a group step: its children run in order. See Group steps. |
assertions | []string | no | — | Expressions evaluated after the response: see Assertions and variables. |
continue_on_error | bool | no | false | A failing step doesn’t abort the test. |
timeout | duration | no | none | Caps each attempt of the step, whatever the protocol. The protocol’s own timeout still applies inside it. A step that hits it errors with step timed out after <timeout>. |
retries | int | no | 0 | Extra attempts when the step errors (it could not run: a transport error, a timeout, a command that could not start). Assertion failures are results and are not retried. Total attempts = 1 + retries; the console shows the count when it is more than one. |
retry_delay | duration | no | 200ms | Constant delay between attempts. |
files | map | no | — | Tabular files (csv/tsv/xlsx) this step produced, loaded for assertions: see CLI steps. |
artifacts | []string | no | — | Files or directories the step produced that are kept with the run: paths or globs, relative to the test file, templated. Copied into the run’s artifacts directory after the step, pass or fail (see Outputs). A pattern that matches nothing is noted on the step. |
tags | []string | no | — | Per-step tags (reporting granularity). |
HTTP steps
Section titled “HTTP steps”tests: - name: create-user steps: - name: create http: url: "{{baseurl}}/users" # required method: POST # default GET headers: Content-Type: application/json Authorization: "Bearer {{token}}" query: # appended to the URL; existing params preserved page: "1" body: type: raw # raw (default) | form | multipart payload: '{"name": "{{username}}"}' timeout: 10s # default 30s follow_redirects: true # default true tls: skip_verify: true # self-signed certs in dev response: type: json # json (default) | text | binary variables: # extract into scope for later steps user_id: "response.id" first_tag: "jq: .response.tags[0]" assertions: - status_code == 201 - response.id != null - response.name == "{{username}}" - duration_ms < 500 - headers["content-type"] contains "application/json"Body types:
-
raw:payloadstring sent as-is; setContent-Typeyourself. -
form:fields:map →application/x-www-form-urlencoded. Scalar values only; nested maps/arrays are rejected. -
multipart: each entry infields:is one part. A scalar value is a text part; an object value selects a file or an explicit text part:- name: uploadhttp:url: "{{baseurl}}/meta/product/{{product}}/file/iam"method: POST # required: upload routes are rarely GETbody:type: multipartfields:file: # the part namefile: "./sample.yaml" # ← the object form is what uploads a filefilename: "sample.yaml" # optional; defaults to the basenamecontent_type: "application/yaml" # optional; defaults to application/octet-streamoverwrite: "true" # scalar → ordinary text partassertions:- status_code == 200A text part with an explicit type is
{ value: "text", content_type: "..." }.Those four,
file,filename,content_type,value, are the only keys the object form accepts, and the loader rejects anything else by name:step "upload-datatape-file": body.fields.file: unknown key "filoname"(valid: file, filename, content_type, value): did you mean "filename"?Strict YAML decoding checks keys everywhere in a suite.
fields:is an open map, so its parts get a check of their own: a misspelt key in a part fails the load, naming the key and the nearest valid one.The nesting is the trap.
fields: { file: "./sample.yaml" }, a bare string, sends the path as text, not the file, and no error is raised because a text part namedfileis perfectly legal. The server then reports no file was uploaded. The file part needs the innerfile:key, as above.Remember
method:too: it defaults to GET, and an upload route usually only accepts POST/PUT, so omitting it produces a server-side error that says nothing about the body.
Behavior notes:
- The body builder sets
Content-Typefirst; yourheaders:override it. - Default
User-Agent: kis-test/<version>unless you set one. - A non-2xx status is not an error: assert on
status_code. Only transport failures (DNS, refused connection, timeout) error the step (and are whatretriesretries). - Wire timings are captured per exchange:
dns_ms,connect_ms,tls_ms,ttfb_ms,transfer_ms,total_ms: persisted with--db, andttfb_ms/total_msare assertable. - Cookies: each test gets one cookie jar shared by all its HTTP steps and
before_each/after_eachhooks, so login → authenticated-call flows work without manual cookie plumbing. tls:takesskip_verify, a client certificate for mutual TLS (client_certandclient_key, PEM files, set together), androot_cas, extra CA PEM files trusted on top of the system store. Paths are relative to the file that declares them. gRPC steps take the same block.
Suite defaults: http:
Section titled “Suite defaults: http:”A suite’s http: block sets defaults for every HTTP request in it and its child suites: http:
steps, the http: triggers of stream steps, and SSE streams.
http: base_url: "{{baseurl}}" headers: Authorization: "Bearer {{token}}" timeout: 10s tls: root_cas: [certs/internal-ca.pem]steps: - name: list-users http: url: /users # sent to {{baseurl}}/users, with the Authorization header - name: anonymous-call http: url: /public headers: Authorization: "" # this request sends no Authorization headerbase_urlis joined to a request URL that starts with/. A full URL is sent as written.headersare added to every request that does not set the same header; names compare case-insensitively, so a step’sauthorizationreplaces the suite’sAuthorization. A step that sets a header to""sends none.timeout,follow_redirectsandtlsapply to a request that leaves them unset.- A child suite’s
http:overrides its parent’s key by key;headersmerge. - Values are templates, rendered with each step’s scope when it runs, so
{{token}}can come from abefore_alllogin.
CLI steps
Section titled “CLI steps”cli: runs a command-line program: the “request” is a binary plus its arguments,
environment, working directory and stdin; the “response” is an exit code plus two output
streams. Everything above that boundary, templating, variable extraction, assertions,
retries, data-driven rows, load phases, works exactly as it does for HTTP.
- name: create-a-tenant cli: command: kis # looked up on PATH args: ["tenant", "create", "--name", "{{tenant}}", "--json"] cwd: "fixtures" # optional; defaults to the SUITE directory env: # layered over the parent environment KIS_CONFIG: "{{config_path}}" stdin: "{{payload}}" # written, then closed timeout: 30s # default 60s response: variables: tenant_id: "jq: .json.id" assertions: - exit_code == 0 - stderr == "" - stdout contains "created" - duration_ms < 500| Key | Type | Default | Meaning |
|---|---|---|---|
command | string | — | The executable. Mutually exclusive with shell. |
args | []string | — | Passed verbatim: no shell, no word splitting, no globbing. |
shell | string | — | A command line run through sh -c. Use for pipelines and redirects. |
cwd | string | the suite directory | Relative paths resolve against the suite. |
env | map | — | Layered over the parent environment. |
clean_env | bool | false | Drop the parent environment entirely (takes PATH with it). |
stdin | string | — | Written to stdin, which is then closed. |
timeout | duration | 60s | Wall clock. On expiry the whole process group is signalled. |
grace | duration | 5s | Between SIGTERM and SIGKILL. |
max_output | bytes | 10485760 | Per-stream capture cap. Beyond it truncated is set. |
args: is a list for a reason. Each element stays one argument no matter what it
renders to, so a data-driven row containing a space, a quote or a semicolon becomes a
single argv entry: there is no shell in the path to re-split it. shell: opts out of
that by design; quoting is then yours.
A nonzero exit is not an error. It is a result to assert on, exactly the way a 500 is:
the step runs, captures everything, and your assertions decide. Only a command that could
not be started, or one killed by its own timeout, errors the step.
Timeouts kill the group, not just the child. A command that backgrounds work, most real CLIs and every wrapper script, would otherwise leak those workers on each timed-out iteration; under load that is thousands of orphans.
What you can assert on
Section titled “What you can assert on”| Path | Meaning |
|---|---|
exit_code | Process exit status. |
stdout / stderr / output | Captured text; output is stdout followed by stderr. |
stdout_lines / stderr_lines | Arrays: stdout_lines.length == 5 counts output. |
json | stdout parsed as JSON. Absent when stdout is not JSON. |
duration_ms | Wall clock for the command. |
cpu_ms / max_rss_mb | From the kernel’s rusage: max_rss_mb < 50 is a real gate. |
signal / timed_out / truncated | How it ended, and whether output was capped. |
command | The rendered command line, quoted for display. |
Extracting variables
Section titled “Extracting variables”A CLI’s output is text far more often than JSON, so response.variables takes four forms:
response: variables: exit: "exit_code" # dotted path over the result id: "jq: .json.id" # jq, for tools with a --json mode version: "regex: version (\\S+)" # capture group 1, else the whole match first: "line: 0" # one line; negative counts from the endA jq: path under json on a tool that does not emit JSON reports the stdout it did
produce, rather than jq’s thirty-character truncation of it.
Data-driven
Section titled “Data-driven”Nothing special is needed: a test’s table: fans it out one row at a time and the row’s
columns are in scope, so args: templates against them.
data: cases: rows: - { word: "alpha", n: 5 } - { word: "two words", n: 9 }
tests: - name: echo-length table: cases steps: - name: measure cli: shell: 'printf %s "$WORD" | wc -c' env: { WORD: "{{word}}" } assertions: - stdout contains "{{n}}"Under load
Section titled “Under load”CLI steps run under kis test load like any other protocol, and duration_ms is the
command’s own wall clock. Process spawn dominates per-iteration cost, so this measures
startup time, which for a CLI is a real SLO, rather than throughput of anything else.
stdin:is written to the command and then closed, which suits a command that reads its whole input at once.
Files produced by a step: files:
Section titled “Files produced by a step: files:”files: declares tabular files (CSV, TSV, XLSX) a step wrote, loads them after the step
runs, and exposes them to assertions under files.<name>. It is a step-level block,
not a CLI one: an HTTP download and a SQL export produce files too.
- name: export cli: command: kis args: ["data", "export", "--out", "out/report.csv"] files: report: path: "out/report.csv" # resolved against the suite directory format: csv # csv | tsv | xlsx: inferred from the extension sheet: "Summary" # xlsx only; default is the first sheet key: ["tenant_id"] # row identity → comparison ignores row ORDER ignore: ["generated_at"] # columns excluded from comparison tolerance: 0.01 # numeric epsilon no_header: false # true when the first row is data assertions: - files.report.row_count == 1500 - files.report.columns includes "amount" - files.report.column.amount.sum == 125000 - files.report.column.tenant_id.distinct == 12 - files.report.rows[0].name == "Alpha" - files.report matches_file "fixtures/expected-report.csv"| Path | Meaning |
|---|---|
files.X.row_count / column_count | Counts, excluding the header. |
files.X.columns | The header, as an array. |
files.X.rows[n].<col> | One cell, addressed like a SQL row. |
files.X.column.<col>.sum / avg / min / max | Numeric aggregates, null when the column holds no numbers. |
files.X.column.<col>.count / empty / distinct / numeric | Cell counts. numeric is how many parsed as numbers. |
files.X.column.<col>.values | Every cell, as an array. |
Numbers are read the way exports actually write them: 1,200.50, $300, (150) for an
accounting negative, 25%. numeric tells you how many cells contributed, so a sum over
1,500 values is distinguishable from a sum over the 3 that happened to parse.
matches_file: comparing two files
Section titled “matches_file: comparing two files”- files.report matches_file "fixtures/expected-report.csv"Byte-equality on two CSVs is nearly useless in a test: row order varies, floats format
differently, and every export carries a timestamp. matches_file compares data under
the declaration’s rules:
key:matches rows by identity, so order does not matter. Without a key rows are compared position by position.ignore:drops columns from the comparison (they stay available to aggregates).- numbers compare as numbers, so
100.50equals100.5and1,200equals1200. tolerance:allows float drift.- the two sides need not be the same format: a CSV compares against an XLSX.
The output is a diff, not a boolean:
out/bad.csv does not match fixtures/expected.csv 1 row(s) missing, keyed: t3 1 unexpected row(s), keyed: t9 row t2, column name: expected "Bravo", got "Bravo-RENAMED"A missing column is reported alone: it makes every row under it differ, and listing those cells would bury the cause. Long diffs are capped at 20 cells and say how many were elided.
A file that cannot be read is a warning, not a failure: a missing output file is
usually the finding, and the step’s own exit_code and stderr assertions say why it is
missing far better than a bare “no such file” that replaces them.
gRPC steps
Section titled “gRPC steps”steps: - name: get-user grpc: address: "{{grpc_host}}:443" # required service: user.v1.UserService # required, fully qualified method: GetUser # required tls: # omit entirely for plaintext skip_verify: true metadata: authorization: "Bearer {{token}}" payload: # request message; snake_case or camelCase field names user_id: "{{user_id}}" timeout: 10s # default 30s response: variables: user_name: "response.user.name" assertions: - status_code == 0 # gRPC status code, 0 = OK - response.user.id == "{{user_id}}"- Server reflection: the step finds the method through the target’s reflection service, so the server exposes it. No generated stubs or proto files are needed.
- Calls are unary.
- A non-OK gRPC status is not an error: it lands in
status_code/status_messagefor assertions. - Payloads are marshaled via protojson; responses use
UseProtoNames+EmitUnpopulated, so assert with snake_case field names and expect zero-valued fields to be present. - A run opens one connection per target (address and TLS setting) and resolves each method by reflection once, then reuses both for every later call, including under load.
Database steps
Section titled “Database steps”The step key is db:. sql: is the original spelling and still loads, so existing suites keep
working; prefer db: in new tests.
Named connections
Section titled “Named connections”Declare connections once at suite level and reference them by name, so a DSN (and its credentials) is written in one place:
databases: erp: driver: postgres connection: "{{db_url}}" # templated at execute time like any config string analytics: driver: clickhouse connection: "clickhouse://user:pass@localhost:9000/appdb"
tests: - name: api-matches-db steps: - name: verify-row db: use: erp # supplies driver + connection query: "SELECT id, name FROM users WHERE id = $1" params: ["{{user_id}}"] assertions: - row_count == 1databases: resolves up the suite tree exactly as data: does, so a parent suite.yaml can
declare one connection every descendant reuses. Names resolve at load time: an unknown name
fails the load and lists the ones that exist, rather than reaching the driver with an empty DSN.
use: and driver:/connection: are mutually exclusive: use: already supplies both.
Reusing connections
Section titled “Reusing connections”By default every step opens a connection and closes it again. A connection that sets
reuse: true is instead kept open for the whole run and shared by every step that names it:
databases: erp: driver: postgres connection: "{{db_url}}" reuse: true max_open_conns: 4 # optional; zero leaves database/sql's default max_idle_conns: 2 conn_max_lifetime: 5mReuse is opt-in on purpose: holding connections open changes what the server sees, so a suite
that asserts on pg_stat_activity, or runs against a tight connection budget, keeps the
open/close behavior by not asking for anything.
The cache is keyed by driver + rendered DSN, so one named connection whose DSN embeds a per-test variable still gets a handle per distinct database. Bounds apply when the handle is first created; redeclaring the same DSN with different bounds does not reconfigure a live pool. Handles are closed when the run finishes.
A reused sqlite :memory: connection persists across steps, which is often what you want; the
default closes it between steps and each step starts empty.
Comparing a numeric column to a templated value needs the template unquoted:
rows[0].id == {{user_id}}compares numbers, whilerows[0].id == "{{user_id}}"compares a number to a string and fails.
Inline connections
Section titled “Inline connections”steps: - name: verify-row db: driver: postgres # postgres|postgresql|pg, mysql, clickhouse, duckdb, sqlite|sqlite3 connection: "postgres://user:pass@localhost:5432/db?sslmode=disable" query: "SELECT id, name FROM users WHERE id = $1" params: ["{{user_id}}"] # positional only timeout: 5s # default 10s response: variables: db_name: "rows[0].name" total: "row_count" assertions: - row_count == 1 - rows[0].name == "{{username}}"
- name: seed db: use: erp execute: "INSERT INTO users (id, name) VALUES ($1, $2)" params: ["{{user_id}}", "{{username}}"] assertions: - rows_affected == 1query(SELECT →rows,row_count,columns) andexecute(DML →rows_affected) are mutually exclusive.- Empty
connectionopens an in-memory DuckDB: handy for pure-SQL smoke tests. - Parameters are positional:
paramsis a list, in the driver’s placeholder order. rowskeeps at most a 1000-row sample;[]bytecolumns coerce to string, timestamps to RFC3339Nano.- DB errors mark the step errored (with the error in evidence) rather than aborting the run.
Streaming steps: ws: and sse:
Section titled “Streaming steps: ws: and sse:”Every other step is request/response. A stream is not: it stays open, messages arrive on their own schedule, and what you want to assert is “the event I care about turned up, in time, and looked right”.
A streaming step has two halves that read as a conversation:
triggers:: the things you do. Named, so a script can fire one.expect:: the things the server says back, in order. Each one selects a message, validates it, and can run a script that fires the next trigger.
steps: - name: order-lifecycle sse: url: "{{baseurl}}/events" timeout: 30s
triggers: - name: create # fires at connect http: method: POST url: "{{baseurl}}/orders" headers: {Content-Type: application/json} body: payload: '{"item":"widget","qty":2}' save_as: created # the POST's own response
- name: cancel # idle until a script fires it when: call http: method: POST url: "{{baseurl}}/orders/{{order_id}}/cancel"
expect: - match: "event._kind == 'order.created'" # SELECT the message timeout: 5s save_as: ev_created assertions: # VALIDATE it: all of them run - event.id == {{created.response.id}} - event.qty == 2 - event.total > 0 - event._offset_ms < 500 script: # ANSWER BACK language: js execute: | stream.trigger("cancel", {order_id: event.id});
- match: "event._kind == 'order.cancelled'" timeout: 5s assertions: - event.id == {{ev_created.id}} assertions: - message_count == 2 # whole-stream factsws: is the same, plus the socket-only fields, send:, subprotocols:,
on_close:, and message as the alias for event.
expect: blocks: select, validate, answer
Section titled “expect: blocks: select, validate, answer”| Field | Meaning |
|---|---|
match | an assertion expression evaluated against each arriving message until one satisfies it. Empty matches any message |
assertions | expressions validating the message that matched. Every one runs |
script | runs after the assertions pass, against that message. Usually stream.trigger(...) |
count | wait for this many matches rather than one; save_as then captures the list, and assertions/script run for each |
timeout | deadline for this wait, measured from when it started waiting, not from the start of the step; falls back to the step’s timeout |
save_as | capture the matching message into the test scope |
match selects; assertions validate. The distinction decides what
happens to a wrong message, and it is the most important thing on this
page:
- A message that fails
matchis someone else’s traffic: recorded and skipped, and the wait continues. Real subscriptions carry heartbeats and unrelated events, and failing on the first one would make them untestable. - A message that matched but fails
assertionsfails the step. The event the test asked for did arrive and it was wrong. Resuming the search would hunt for a second one that may never come and report a timeout: naming the wrong problem entirely:
expect #1 matched [1 @ 3ms] order.created but it failed validation: event.total > 0 — expected > 0, got 0What a message looks like
Section titled “What a message looks like”A JSON object’s fields sit at the top level, so message.type and
event.total read naturally. The envelope is always present under
underscore-prefixed names, whatever the payload was:
| Field | Meaning |
|---|---|
_seq | 1-based arrival order |
_offset_ms | milliseconds since the connection opened |
_kind | text/binary for websocket; the event name for SSE |
_id | the SSE event id |
_raw | the payload as received |
data | the parsed payload, or the raw string when it is not JSON |
_offset_ms is why arrival is recorded at all: “the event arrived” is
rarely the question: “it arrived within 200ms of the one before” is.
message and event are the same thing under two names; use whichever
suits the protocol you are thinking about. Both match: and
assertions: are Liquid-rendered first, so they can refer to anything an
earlier step, trigger or wait produced.
Step-level assertions
Section titled “Step-level assertions”After the waits resolve, the step’s own assertions: see
message_count (everything that arrived, including messages past the
retention bound), messages, connected, status_code for SSE, and
every name captured by a save_as. Use them for facts about the stream
as a whole; use a wait’s assertions: for facts about one message.
Failure, and why the distinction matters
Section titled “Failure, and why the distinction matters”- A wait that is never satisfied fails the step.
- A wait whose assertions fail fails the step.
- A connection that could not be opened errors it: the question was never asked.
An unmet wait reports what it wanted and what arrived instead, because a bare “timeout” on a busy stream tells you nothing:
expect #1 timed out after 300ms waiting for "event._kind == 'order.created'" — saw instead: [1 @ 5ms] heartbeat {"n":1} [2 @ 11ms] heartbeat {"n":2}triggers: blocks: the things you do
Section titled “triggers: blocks: the things you do”A trigger is http: or script:, and fires either at connect or when a
script calls it.
| Field | Meaning |
|---|---|
name | identifies it for stream.trigger(name, params). Required for when: call |
when | connect (default) fires in order once the stream is open; call sits idle until a script fires it |
http | a request: the same block an http: step takes, so auth, bodies, uploads and templating all work |
script | a script instead, with the same task namespaces a script: step has. Mutually exclusive with http |
save_as | captures the outcome: for http, {status_code, response, headers, duration_ms}; for script, whatever it returned |
repeat | fire it this many times in sequence. save_as then captures the list |
delay | wait this long before firing. Rarely needed; reach for it when a server registers the subscription asynchronously after the handshake |
trigger: (singular) is shorthand for a one-element triggers:. Setting
both is a load error rather than a silent choice between them.
when: connect triggers fire in order, and each sees what the ones
before it captured, so the second can address what the first created. A
failure stops the sequence: amending an order that was never created just
produces a second, more confusing error on top of the real one. The step
reports trigger "amend" (#2): ... rather than the timeout it caused.
To fire the same trigger N times, use repeat:, not N copies of the
block:
triggers: - name: create repeat: 3 http: {method: POST, url: "{{baseurl}}/orders"} save_as: made # a list of 3 expect: - match: "event._kind == 'order.created'" count: 3 # order-independent, unlike three waits assertions: - event.total > 0 # runs against each of the threestream.trigger(): answering back
Section titled “stream.trigger(): answering back”Inside an expect: script, stream is in scope:
stream.trigger("cancel", {order_id: event.id});The params become variables when that trigger’s config is templated, so
the trigger above can write {{order_id}}. It returns
{success: true, fired: "cancel"}, or {success: false, error: ...}
naming the triggers that do exist if you mistyped the name: a silently
ignored call would surface as an unexplained timeout much later.
It returns immediately. The request goes out on its own goroutine, so
the reader is not stalled and the offsets of the messages it provokes
still measure the server rather than the runner. There is no synchronous
variant: the point of a stream step is that the response comes back as a
message, which is what the next expect: is for.
The script itself runs synchronously on the reader, so keep it short: a
slow script delays later messages and inflates their offsets. It can also
test.set_variable() into the test scope, and test.fail() to fail the
step.
Timing, and the race this closes
Section titled “Timing, and the race this closes”Correlating against a trigger’s capture is safe even when the push wins
the race. A service that publishes inside the same transaction that
answers the request can deliver the event to a subscriber before the
runner has finished reading the POST’s own response. A match: or an
assertions: entry that reads a variable the triggers have not set holds its
message until the triggers have settled, rather than templating
{{created.response.id}} against nothing. A wait whose variables are all
already resolved is judged the instant its message lands.
A non-2xx trigger does not fail the step: “POST returns 409 and the service pushes a conflict event” is a real test. But because an unexplained timeout is so often a trigger that did not do what the author assumed, the status is carried into the timeout message:
expect #1 timed out after 600ms waiting for "event._kind == 'order.created'" — no messages arrived while waiting — trigger #1 POST http://localhost:8080/orders returned 422 Unprocessable EntityA trigger that could not be sent at all is different, as is a script
trigger that calls test.fail(): the step errors immediately rather than
waiting out a deadline for a push that provably cannot come.
Other fields
Section titled “Other fields”max_messages (ws) / max_events (sse) bound what is retained as
evidence, default 1000; the count stays accurate beyond it. ws: also
takes subprotocols:, send:, on_close: and tls:; sse: takes
method:, query:, body: and tls: like an HTTP step.
Neither client reconnects. A silent reconnect would hide exactly the disconnection a test may exist to catch.
Script steps
Section titled “Script steps”steps: - name: custom-check script: language: js # default js. Also: js:v8, lua, starlark, cel, expr execute: | const r = response; // parsed body of the previous HTTP step if (!r.items.every(i => i.price > 0)) { test.fail("non-positive price found", 400); } test.set_variable("item_count", r.items.length); test.success("all items valid", 200); timeout: 10s # default 5s params: { max: 10 } # extra vars merged over the scopeexecute (inline) and file: (path relative to the test YAML’s directory) are mutually
exclusive. See Scripting for the full test.* API, injected variables, and
pass/fail semantics.
Browser steps: playwright:
Section titled “Browser steps: playwright:”A playwright: step drives a browser with Playwright, under Node.
It runs in the same test as HTTP, database and CLI steps, so one test can create data through
the API, check it in the UI, and confirm it in the database.
Playwright comes from the project: the step runs Node in its working directory, and
playwright or @playwright/test must resolve from there. The browser is
Obscura unless you choose another: a lightweight
headless browser that Playwright drives over the Chrome DevTools Protocol (CDP). Install
Playwright in the project, and Obscura on the machine that runs the tests, with obscura and
obscura-worker side by side on PATH (or TALOS_OBSCURA naming obscura):
npm install -D @playwright/testcurl -L https://github.com/h4ckf0r0day/obscura/releases/download/v0.2.3/obscura-x86_64-linux.tar.gz | tar xz -C ~/.local/binFor a browser Playwright launches itself, install it with npx playwright install chromium
(or firefox, webkit) and choose it as below.
Before a run starts, the runner checks that Node is on PATH, that Playwright resolves, and
that Obscura is installed, for every browser step the run will reach, and stops with what to do
when one is missing.
Choosing the browser
Section titled “Choosing the browser”| Browser | How it runs |
|---|---|
obscura (the default) | The runner starts obscura serve for the step on 127.0.0.1, with loopback and private-network addresses allowed so it reaches the service under test, connects Playwright to it over CDP, and stops it when the step ends. |
chromium, firefox, webkit | Playwright launches its own browser. |
cdp: an endpoint | Connects to a browser that is already running and starts nothing: a shared Obscura server (obscura serve --port 9222 --allow-private-network serves ws://127.0.0.1:9222/devtools/browser), or a Chrome started with --remote-debugging-port. |
The most specific choice wins:
- the step’s
browser:orcdp:; - the nearest suite’s
playwright:block, which applies to every browser step in the suite and the suites below it; - the run’s
--browser(default$TALOS_BROWSER); obscura.
playwright: browser: chromium # or: cdp: ws://browsers.internal:9222A spec step runs the browsers of the project’s playwright.config unless one of these chooses
obscura or a cdp: endpoint. The runner then hands the endpoint to the specs as
TALOS_CDP_ENDPOINT, and the specs use it through a fixture they import test and expect
from:
const base = require('@playwright/test');exports.test = base.test.extend({ browser: [async ({ playwright }, use) => { const endpoint = process.env.TALOS_CDP_ENDPOINT; const browser = endpoint ? await playwright.chromium.connectOverCDP(endpoint) : await playwright.chromium.launch(); await use(browser); await browser.close(); }, { scope: 'worker' }],});exports.expect = base.expect;On Obscura, page.setContent() in a script step waits for the content to be committed
(waitUntil: 'commit'), because Obscura reports no load event for a page filled in place; in
a spec, pass { waitUntil: 'commit' } or open a data: URL instead. The browser state an
earlier step left carries into the next script step as cookies and local storage, and into a
spec run as cookies.
A step takes one of two forms.
Script form
Section titled “Script form”script: (or file:) is a JavaScript body run in a fresh browser context. await works at the
top level.
steps: - name: login-via-api http: method: POST url: "{{app_url}}/api/login" headers: { Content-Type: application/json } body: payload: '{"email": "{{email}}", "password": "{{password}}"}'
- name: dashboard-shows-the-user playwright: base_url: "{{app_url}}" script: | await page.goto("/dashboard"); await expect(page.getByTestId("user-name")).toHaveText(vars.username); talos.set("order_count", await page.getByTestId("order").count()); talos.log("dashboard at", page.url()); assertions: - passed == 1 - console_errors.length == 0
- name: orders-match db: use: app query: "SELECT count(*) AS n FROM orders WHERE user_email = '{{email}}'" assertions: - rows[0].n == {{order_count}}| In scope | What it is |
|---|---|
page, context, browser | Playwright’s page, browser context and browser |
expect | @playwright/test’s expect with auto-waiting matchers (present when @playwright/test is installed) |
vars | the test’s variables |
talos.set(name, value) | set a variable for the steps after this one; export: publishes it to suite or run memory |
talos.log(...) | a line printed under the step and kept in JUnit system-out |
require | Node’s require, resolving from the step’s working directory |
A thrown error or a failing expect fails the step with its message. The step errors instead
when it cannot run at all: Node missing, Playwright not installed, a browser that does not
launch.
Spec form
Section titled “Spec form”spec: runs existing Playwright Test specs, the same as npx playwright test, with the
project’s playwright.config applying. The runner reads the result from Playwright’s JSON
report.
steps: - name: checkout-journey playwright: spec: e2e/checkout.spec.ts # file, directory or pattern; `.` runs every spec grep: "@smoke" project: chromium base_url: "{{app_url}}" assertions: - failed == 0 - flaky == 0A spec reaches the test’s variables and hands values back through environment variables:
| Variable | What it is |
|---|---|
TALOS_VARS | path of a JSON file holding the test’s variables |
TALOS_EXPORT | path to write a JSON object to; its keys become variables for the steps after this one |
TALOS_BASE_URL | the step’s base_url |
TALOS_STORAGE_STATE | path of a storage state holding the cookies the test’s HTTP steps collected for base_url and the session an earlier browser step left; set when there is one |
use: { baseURL: process.env.TALOS_BASE_URL, storageState: process.env.TALOS_STORAGE_STATE,},Fields
Section titled “Fields”| Key | Form | Default | Meaning |
|---|---|---|---|
script / file | script | — | The body, inline or from a file relative to the test file. |
spec | spec | — | Spec files to run, as playwright test takes them. |
grep / project / config | spec | — | --grep, --project, --config. |
browser | both | obscura | obscura, chromium, firefox or webkit; see Choosing the browser. A spec step takes only obscura. |
cdp | both | — | The DevTools endpoint of a browser that is already running, to connect to instead. |
headless | both | true | false shows the browser (--headed in spec form). |
base_url | both | — | Resolves relative page.goto() paths, and is the origin whose cookies are carried in from HTTP steps. |
viewport | script | Playwright’s | { width, height }. |
trace | both | on_failure | always, on_failure or never. In spec form, --trace on, retain-on-failure or off. |
screenshot | script | on_failure | A full-page screenshot when the script ends. |
video | script | never | A video of the session. |
cwd | both | the test file’s directory | Where Node runs and node_modules resolves from. |
env | both | — | Extra environment variables, templated. |
node | both | node | The Node.js binary. |
timeout | both | 5m | Caps the whole run. |
export | both | — | Publish the variables the step set to suite or run memory. |
Keys that apply to the other form are load errors: in spec form, the viewport, screenshots,
video and any browser other than Obscura come from playwright.config.
One session across steps
Section titled “One session across steps”The steps of a test share one session:
- The HTTP steps’ cookies for
base_urlare added to the browser, so an API login carries into the UI. - The browser’s cookies go back into the test’s cookie jar afterwards, so an HTTP step after a UI login is logged in too.
- A script step’s browser state (cookies and local storage) carries to the test’s next browser step, so each step starts where the last one left off. A retried test starts a fresh session.
What you can assert on
Section titled “What you can assert on”| Field | What it is |
|---|---|
status | passed or failed |
passed / failed / skipped / flaky | test counts (script form: one test) |
tests | each test: title, file, project, status, duration_ms, retries, error |
errors | failure messages, and errors Playwright reported outside a test |
console_errors / page_errors | what the page logged with console.error, and uncaught exceptions (script form) |
url / title | the page’s when the script finished (script form) |
Variables the step set are in scope too.
Evidence
Section titled “Evidence”The report, traces, screenshots and videos are kept in the run’s artifacts directory (see
Outputs). Open a trace with npx playwright show-trace <path>. The step’s
result is stored in step_evidence with --db.
kis test load runs a suite’s other step types. A browser per virtual user would measure the
load generator rather than the service, so a load profile that reaches a browser step stops
before it starts and names the step; exclude browser tests from load runs with
--exclude-tags.
Hooks are steps that run at suite boundaries; they use the identical step schema.
| Hook | Runs | Failure behavior |
|---|---|---|
before_all | once per suite, before its tests | the suite’s tests do not run; each is reported errored, naming the failed step. Child suites and after_all still run |
before_each | full suite chain root → leaf, before every test (and every table row) | test marked errored, steps skipped; after_each still runs |
after_each | leaf → root, after every test, pass or fail | logged, does not change the test result |
after_all | once per suite, always: even after errors | logged, does not change the suite result |
Variables set by before_all persist in the suite scope (visible to all tests in the suite and
to child suites). Variables set by before_each live in the test scope. This is how auth
tokens propagate:
before_each: - name: authenticate http: url: "{{baseurl}}/auth" method: POST body: payload: '{"client_id": "{{client_id}}"}' response: variables: token: "response.access_token"Suite-level hooks (before_all/after_all) get their own cookie jar; per-test hooks share the
test’s jar.
Hooks are declared in suite.yaml, or at the top of a file run on its own with -t file.yaml.
A hook in any other file of a directory suite fails the load; see
Where suite-level keys live.
Group steps
Section titled “Group steps”A step with steps: instead of a protocol is a group: its children run in order, in the test’s
scope, so a variable one child extracts is visible to the next and to the steps after the
group.
steps: - name: login retries: 2 timeout: 20s steps: - name: get-token http: method: POST url: /auth/token response: variables: token: "response.access_token" - name: check-token http: url: /auth/verify headers: { Authorization: "Bearer {{token}}" } assertions: - status_code == 200 assertions: - token != null
- name: create-order http: method: POST url: /orders headers: { Authorization: "Bearer {{token}}" }- The group passes when every child passes. The first failing child ends the group unless the
child sets
continue_on_error; the group then fails, and its error names the child. timeoutcaps the whole group,retriesruns the whole group again when it errors, andcontinue_on_erroron the group lets the test go on past it.assertions:on a group run after its children pass, against the scope they left.- Children can be groups. The console prints a group’s children indented beneath it, the
group’s line says how many passed (
2/2 steps passed), and--dbstores each child instep_runswith the group as itsparent_step_id.
Reusing steps
Section titled “Reusing steps”A step block used by many tests lives in a shared file and is pulled in with import:. Keep it
in a _-prefixed directory so the loader does not treat it as a suite.
steps: - name: login steps: - name: get-token http: { method: POST, url: /auth/token, response: { variables: { token: "response.access_token" } } } - name: check-token http: { url: /auth/verify, headers: { Authorization: "Bearer {{token}}" } } - name: health http: { url: /health }import: shared: ../_shared/auth.yamltests: - name: checkout steps: - insert: $shared:/steps[name=login] # the login group, as one step - splice: $shared:/steps # or every step in the file, in order - name: pay http: { method: POST, url: /pay }insert: and splice: both place the selected steps into the list; a group arrives as one
step with its children. Paths are relative to the importing file. Hooks take shared steps the
same way.
Tags and filtering
Section titled “Tags and filtering”kis test run -t tests/ --tags smoke # tagged smokekis test run -t tests/ --tags "smoke,api" # smoke AND apikis test run -t tests/ --tags smoke --tags nightly # smoke OR nightlykis test run -t tests/ --tags api --exclude-tags wipCommas inside one --tags value = AND; repeating the flag = OR; --exclude-tags uses the same
syntax and wins ties. A test’s effective tags are the union of its own tags and every ancestor
suite’s tags.
skip: and only:
Section titled “skip: and only:”tests: - name: refund-to-card skip: "the sandbox has no card processor (PAY-12)" steps: [...]
- name: new-checkout-flow only: true # while writing it: run just this one steps: [...]skip: trueorskip: "<reason>"on a test, or in asuite.yamlfor the whole suite, reports the test as skipped with the reason. The console shows it, and JUnit carries it as<skipped message="...">.- When any test in the run sets
only: true, only those tests run; the console says how many.--forbid-onlymakes such a run fail instead, which is the setting for CI. - A suite’s hooks run only when at least one test in the suite or below it will run. Narrowing a
run with
--tags,only:orskip:does not set up suites none of whose tests run. kis test loadleaves skipped tests out of the workload and honoursonly:the same way.
Validation
Section titled “Validation”--dry-run loads the suite tree, validates it, and renders every step’s templates without
running any step. Each test reports as skipped; a step whose template does not render reports
as errored.
What a step would set when it runs stands in as a placeholder, <name>, for the steps after
it: its response.variables, a stream’s save_as, the names a script passes to
set_variable, set_suite_var or set_run_var, and the names a browser script passes to
talos.set, published wherever the step’s export: or setter would put them. A dry run
therefore shows each request as it would be sent, with <order_id> where the extracted id
goes, and an undefined variable it reports is one nothing sets: a misspelling, or a value the
environment does not provide. A name a script computes at run time cannot be known in advance.
A step declares exactly one protocol. A step that declares two (http: and cli:, say) is a
load error naming both.
Unknown keys are errors
Section titled “Unknown keys are errors”Loading is strict: a key the schema does not know fails the load.
This matters more than it sounds. The classic mistake is indenting assertions: inside the
http: block instead of beside it. Under a lenient parser the key is silently dropped, the step
has nothing to assert, and every run reports a pass, including with a deliberately wrong
status_code. A test that quietly asserts nothing is worse than one that fails.
# Wrong: assertions belongs to the step, not the http config- name: ready-check http: url: "{{baseurl}}/ready" assertions: - status_code == 200line 9: field assertions not found in type types.HTTPConfig "assertions" belongs at step level — unindent it to line up with "http:", not inside itWhere a misplaced key is valid at an outer level the error says so; otherwise it lists the keys
that are valid where you wrote it. The same strictness catches export: targets other than
suite/run, a use: naming a connection that is not declared, a step that sets both
db: and sql:, and a suite-scoped key (variables:, data:, databases:, a hook) written
in a test file of a directory suite, where it would otherwise never apply.
Everything else load-time validation enforces: every test has a name, every step has a name
and exactly one protocol key or steps:, hooks the same way, and every database step resolves to a driver
plus a connection. On a cli: step it also enforces command: XOR shell:, and rejects
args: beside shell:: sh -c ignores anything appended, so those arguments would
simply never be passed.