Skip to content
Talk to our solutions team

Test Definitions

This is the schema for the v2 runner (kis test run). Tests are YAML, organized in directories, with three levels:

Suite (directory or file)
└── Test (unit of pass/fail)
└── Step (one protocol action + assertions)

Suites in the older steps:/testcases:/scenarios:/testplans: format run with kis test tests; see the CLI Reference.

kis test run -t <path> accepts a single YAML file (one flat suite) or a directory:

  • Every directory is a suite; subdirectories become child suites, recursively.
  • A file named suite.yaml in a directory supplies suite-level config (name, variables, tags, data, databases, hooks, parallel). All other *.yaml/*.yml files contribute only their tests:. A test file that declares variables:, data:, databases:, or a hook at file level fails the load, with an error naming the file, the key, and the directory’s suite.yaml. See Where suite-level keys live.
  • Files and subdirectories are processed in sorted order.
  • Skipped: hidden entries (leading .), directories starting with _ (the convention for shared YAML pulled in via import:), node_modules and vendor (package installs, such as a browser suite’s Playwright), and directories containing no YAML anywhere.
  • Multiple YAML files in one directory contribute their tests to the same suite.
tests/
├── suite.yaml # root suite config: variables, tags, hooks
├── env.yaml # env file (auto-discovered: not a test file)
├── users/
│ ├── suite.yaml # child suite; inherits root variables/tags
│ ├── crud.yaml # tests
│ └── permissions.yaml # tests
└── orders/
└── checkout.yaml # child suite without its own suite.yaml

Files that declare a top-level import: block are expanded through yamlplus (merge: / replace: / insert: / splice: verbs with selectors), with the file’s directory as the base for relative references. Files without import: are parsed as plain YAML. Put shared step blocks in a _-prefixed directory so they aren’t loaded as suites themselves. See Reusing steps for the common case.

Top-level keys of suite.yaml (or of a single-file suite):

KeyTypeDefaultMeaning
namestringdirectory/file basenameSuite identifier.
tags[]string—Accumulate down the tree: effective tags = ancestor suite tags ∪ test tags.
variablesmap—Visible to every test in this suite and children; deeper suite overrides shallower.
datamap[name]DataSource—Named data sources for table: iteration: see Data-driven tests.
skipbool or string—Skips every test in this suite and its child suites, with the reason; their hooks do not run.
httpHTTPDefaults—Defaults for every HTTP request in this suite and its children: base_url, headers, timeout, follow_redirects, tls. See Suite defaults.
parallelboolfalseRun this suite’s tests concurrently.
max_parallelint0 = unboundedConcurrency cap when parallel: true.
tests[]Test—Tests declared in this file.
before_all / before_each / after_each / after_all[]Step—Hooks; same shape as test steps. See below.

variables, data, databases, http, skip, and the four hooks are suite-scoped. Where they are read depends on how the run is started:

Entry pointWhere suite-scoped keys are read
-t dir/Only from each directory’s suite.yaml.
-t file.yamlFrom the file itself, which is the whole suite.

In a directory run, a test file (any YAML file other than suite.yaml) that declares one of these keys at file level fails the load:

load tests/products.yaml: suite-level keys in a test file are not applied when loading a directory; only suite.yaml is read for them:
before_all: move it to tests/suite.yaml, or into the test's own steps:

Move the key to the directory’s suite.yaml. A value that only one test needs can go on the test instead: variables: on the test, or a setup step at the start of its steps:. A file that has to run the same way under both entry points keeps its suite-scoped keys in suite.yaml and only tests: in the file.

KeyTypeRequiredDefaultMeaning
namestringyes—Unique within the suite.
tags[]stringno—Combined with suite-chain tags for --tags filtering.
skipbool or stringno—true, or the reason as a string. The test reports as skipped, with the reason, and nothing in it runs.
onlyboolnofalseWhen any test in the run sets it, only those tests run. See Tags and filtering.
variablesmapno—Merged over the suite snapshot before hooks/steps run.
tablestringno—Data source name; runs the steps once per row.
steps[]Stepyes—Executed in order.
continue_on_errorboolnofalseKeep running steps after a failure (reports all failures); also lets table iteration continue past a failing row.
timeoutdurationnononeCaps each attempt of the test: its before_each chain and its steps. A test that outlasts it errors with test timed out after <timeout>; after_each still runs.
retriesintno0Runs a test that failed or errored again from the start, with a fresh scope and cookie jar, up to this many more times. The result is the last attempt’s; the console prints ↻ <test> passed on attempt N so a flaky test stays visible, and --db stores the count in test_runs.retry_count.

Each step declares exactly one protocol key.

KeyTypeRequiredDefaultMeaning
namestringyes—Step identifier.
http / grpc / db / script / cli / ws / sse / playwrightobjectone, or steps—Protocol config (below). sql: is accepted as the old name for db:.
steps[]Stepinstead of a protocol—Makes this a group step: its children run in order. See Group steps.
assertions[]stringno—Expressions evaluated after the response: see Assertions and variables.
continue_on_errorboolnofalseA failing step doesn’t abort the test.
timeoutdurationnononeCaps each attempt of the step, whatever the protocol. The protocol’s own timeout still applies inside it. A step that hits it errors with step timed out after <timeout>.
retriesintno0Extra attempts when the step errors (it could not run: a transport error, a timeout, a command that could not start). Assertion failures are results and are not retried. Total attempts = 1 + retries; the console shows the count when it is more than one.
retry_delaydurationno200msConstant delay between attempts.
filesmapno—Tabular files (csv/tsv/xlsx) this step produced, loaded for assertions: see CLI steps.
artifacts[]stringno—Files or directories the step produced that are kept with the run: paths or globs, relative to the test file, templated. Copied into the run’s artifacts directory after the step, pass or fail (see Outputs). A pattern that matches nothing is noted on the step.
tags[]stringno—Per-step tags (reporting granularity).
tests:
- name: create-user
steps:
- name: create
http:
url: "{{baseurl}}/users" # required
method: POST # default GET
headers:
Content-Type: application/json
Authorization: "Bearer {{token}}"
query: # appended to the URL; existing params preserved
page: "1"
body:
type: raw # raw (default) | form | multipart
payload: '{"name": "{{username}}"}'
timeout: 10s # default 30s
follow_redirects: true # default true
tls:
skip_verify: true # self-signed certs in dev
response:
type: json # json (default) | text | binary
variables: # extract into scope for later steps
user_id: "response.id"
first_tag: "jq: .response.tags[0]"
assertions:
- status_code == 201
- response.id != null
- response.name == "{{username}}"
- duration_ms < 500
- headers["content-type"] contains "application/json"

Body types:

  • raw: payload string sent as-is; set Content-Type yourself.

  • form: fields: map → application/x-www-form-urlencoded. Scalar values only; nested maps/arrays are rejected.

  • multipart: each entry in fields: is one part. A scalar value is a text part; an object value selects a file or an explicit text part:

    - name: upload
    http:
    url: "{{baseurl}}/meta/product/{{product}}/file/iam"
    method: POST # required: upload routes are rarely GET
    body:
    type: multipart
    fields:
    file: # the part name
    file: "./sample.yaml" # ← the object form is what uploads a file
    filename: "sample.yaml" # optional; defaults to the basename
    content_type: "application/yaml" # optional; defaults to application/octet-stream
    overwrite: "true" # scalar → ordinary text part
    assertions:
    - status_code == 200

    A text part with an explicit type is { value: "text", content_type: "..." }.

    Those four, file, filename, content_type, value, are the only keys the object form accepts, and the loader rejects anything else by name:

    step "upload-datatape-file": body.fields.file: unknown key "filoname"
    (valid: file, filename, content_type, value): did you mean "filename"?

    Strict YAML decoding checks keys everywhere in a suite. fields: is an open map, so its parts get a check of their own: a misspelt key in a part fails the load, naming the key and the nearest valid one.

    The nesting is the trap. fields: { file: "./sample.yaml" }, a bare string, sends the path as text, not the file, and no error is raised because a text part named file is perfectly legal. The server then reports no file was uploaded. The file part needs the inner file: key, as above.

    Remember method: too: it defaults to GET, and an upload route usually only accepts POST/PUT, so omitting it produces a server-side error that says nothing about the body.

Behavior notes:

  • The body builder sets Content-Type first; your headers: override it.
  • Default User-Agent: kis-test/<version> unless you set one.
  • A non-2xx status is not an error: assert on status_code. Only transport failures (DNS, refused connection, timeout) error the step (and are what retries retries).
  • Wire timings are captured per exchange: dns_ms, connect_ms, tls_ms, ttfb_ms, transfer_ms, total_ms: persisted with --db, and ttfb_ms/total_ms are assertable.
  • Cookies: each test gets one cookie jar shared by all its HTTP steps and before_each/ after_each hooks, so login → authenticated-call flows work without manual cookie plumbing.
  • tls: takes skip_verify, a client certificate for mutual TLS (client_cert and client_key, PEM files, set together), and root_cas, extra CA PEM files trusted on top of the system store. Paths are relative to the file that declares them. gRPC steps take the same block.

A suite’s http: block sets defaults for every HTTP request in it and its child suites: http: steps, the http: triggers of stream steps, and SSE streams.

suite.yaml
http:
base_url: "{{baseurl}}"
headers:
Authorization: "Bearer {{token}}"
timeout: 10s
tls:
root_cas: [certs/internal-ca.pem]
steps:
- name: list-users
http:
url: /users # sent to {{baseurl}}/users, with the Authorization header
- name: anonymous-call
http:
url: /public
headers:
Authorization: "" # this request sends no Authorization header
  • base_url is joined to a request URL that starts with /. A full URL is sent as written.
  • headers are added to every request that does not set the same header; names compare case-insensitively, so a step’s authorization replaces the suite’s Authorization. A step that sets a header to "" sends none.
  • timeout, follow_redirects and tls apply to a request that leaves them unset.
  • A child suite’s http: overrides its parent’s key by key; headers merge.
  • Values are templates, rendered with each step’s scope when it runs, so {{token}} can come from a before_all login.

cli: runs a command-line program: the “request” is a binary plus its arguments, environment, working directory and stdin; the “response” is an exit code plus two output streams. Everything above that boundary, templating, variable extraction, assertions, retries, data-driven rows, load phases, works exactly as it does for HTTP.

- name: create-a-tenant
cli:
command: kis # looked up on PATH
args: ["tenant", "create", "--name", "{{tenant}}", "--json"]
cwd: "fixtures" # optional; defaults to the SUITE directory
env: # layered over the parent environment
KIS_CONFIG: "{{config_path}}"
stdin: "{{payload}}" # written, then closed
timeout: 30s # default 60s
response:
variables:
tenant_id: "jq: .json.id"
assertions:
- exit_code == 0
- stderr == ""
- stdout contains "created"
- duration_ms < 500
KeyTypeDefaultMeaning
commandstring—The executable. Mutually exclusive with shell.
args[]string—Passed verbatim: no shell, no word splitting, no globbing.
shellstring—A command line run through sh -c. Use for pipelines and redirects.
cwdstringthe suite directoryRelative paths resolve against the suite.
envmap—Layered over the parent environment.
clean_envboolfalseDrop the parent environment entirely (takes PATH with it).
stdinstring—Written to stdin, which is then closed.
timeoutduration60sWall clock. On expiry the whole process group is signalled.
graceduration5sBetween SIGTERM and SIGKILL.
max_outputbytes10485760Per-stream capture cap. Beyond it truncated is set.

args: is a list for a reason. Each element stays one argument no matter what it renders to, so a data-driven row containing a space, a quote or a semicolon becomes a single argv entry: there is no shell in the path to re-split it. shell: opts out of that by design; quoting is then yours.

A nonzero exit is not an error. It is a result to assert on, exactly the way a 500 is: the step runs, captures everything, and your assertions decide. Only a command that could not be started, or one killed by its own timeout, errors the step.

Timeouts kill the group, not just the child. A command that backgrounds work, most real CLIs and every wrapper script, would otherwise leak those workers on each timed-out iteration; under load that is thousands of orphans.

PathMeaning
exit_codeProcess exit status.
stdout / stderr / outputCaptured text; output is stdout followed by stderr.
stdout_lines / stderr_linesArrays: stdout_lines.length == 5 counts output.
jsonstdout parsed as JSON. Absent when stdout is not JSON.
duration_msWall clock for the command.
cpu_ms / max_rss_mbFrom the kernel’s rusage: max_rss_mb < 50 is a real gate.
signal / timed_out / truncatedHow it ended, and whether output was capped.
commandThe rendered command line, quoted for display.

A CLI’s output is text far more often than JSON, so response.variables takes four forms:

response:
variables:
exit: "exit_code" # dotted path over the result
id: "jq: .json.id" # jq, for tools with a --json mode
version: "regex: version (\\S+)" # capture group 1, else the whole match
first: "line: 0" # one line; negative counts from the end

A jq: path under json on a tool that does not emit JSON reports the stdout it did produce, rather than jq’s thirty-character truncation of it.

Nothing special is needed: a test’s table: fans it out one row at a time and the row’s columns are in scope, so args: templates against them.

data:
cases:
rows:
- { word: "alpha", n: 5 }
- { word: "two words", n: 9 }
tests:
- name: echo-length
table: cases
steps:
- name: measure
cli:
shell: 'printf %s "$WORD" | wc -c'
env: { WORD: "{{word}}" }
assertions:
- stdout contains "{{n}}"

CLI steps run under kis test load like any other protocol, and duration_ms is the command’s own wall clock. Process spawn dominates per-iteration cost, so this measures startup time, which for a CLI is a real SLO, rather than throughput of anything else.

stdin: is written to the command and then closed, which suits a command that reads its whole input at once.

files: declares tabular files (CSV, TSV, XLSX) a step wrote, loads them after the step runs, and exposes them to assertions under files.<name>. It is a step-level block, not a CLI one: an HTTP download and a SQL export produce files too.

- name: export
cli:
command: kis
args: ["data", "export", "--out", "out/report.csv"]
files:
report:
path: "out/report.csv" # resolved against the suite directory
format: csv # csv | tsv | xlsx: inferred from the extension
sheet: "Summary" # xlsx only; default is the first sheet
key: ["tenant_id"] # row identity → comparison ignores row ORDER
ignore: ["generated_at"] # columns excluded from comparison
tolerance: 0.01 # numeric epsilon
no_header: false # true when the first row is data
assertions:
- files.report.row_count == 1500
- files.report.columns includes "amount"
- files.report.column.amount.sum == 125000
- files.report.column.tenant_id.distinct == 12
- files.report.rows[0].name == "Alpha"
- files.report matches_file "fixtures/expected-report.csv"
PathMeaning
files.X.row_count / column_countCounts, excluding the header.
files.X.columnsThe header, as an array.
files.X.rows[n].<col>One cell, addressed like a SQL row.
files.X.column.<col>.sum / avg / min / maxNumeric aggregates, null when the column holds no numbers.
files.X.column.<col>.count / empty / distinct / numericCell counts. numeric is how many parsed as numbers.
files.X.column.<col>.valuesEvery cell, as an array.

Numbers are read the way exports actually write them: 1,200.50, $300, (150) for an accounting negative, 25%. numeric tells you how many cells contributed, so a sum over 1,500 values is distinguishable from a sum over the 3 that happened to parse.

- files.report matches_file "fixtures/expected-report.csv"

Byte-equality on two CSVs is nearly useless in a test: row order varies, floats format differently, and every export carries a timestamp. matches_file compares data under the declaration’s rules:

  • key: matches rows by identity, so order does not matter. Without a key rows are compared position by position.
  • ignore: drops columns from the comparison (they stay available to aggregates).
  • numbers compare as numbers, so 100.50 equals 100.5 and 1,200 equals 1200.
  • tolerance: allows float drift.
  • the two sides need not be the same format: a CSV compares against an XLSX.

The output is a diff, not a boolean:

out/bad.csv does not match fixtures/expected.csv
1 row(s) missing, keyed: t3
1 unexpected row(s), keyed: t9
row t2, column name: expected "Bravo", got "Bravo-RENAMED"

A missing column is reported alone: it makes every row under it differ, and listing those cells would bury the cause. Long diffs are capped at 20 cells and say how many were elided.

A file that cannot be read is a warning, not a failure: a missing output file is usually the finding, and the step’s own exit_code and stderr assertions say why it is missing far better than a bare “no such file” that replaces them.

steps:
- name: get-user
grpc:
address: "{{grpc_host}}:443" # required
service: user.v1.UserService # required, fully qualified
method: GetUser # required
tls: # omit entirely for plaintext
skip_verify: true
metadata:
authorization: "Bearer {{token}}"
payload: # request message; snake_case or camelCase field names
user_id: "{{user_id}}"
timeout: 10s # default 30s
response:
variables:
user_name: "response.user.name"
assertions:
- status_code == 0 # gRPC status code, 0 = OK
- response.user.id == "{{user_id}}"
  • Server reflection: the step finds the method through the target’s reflection service, so the server exposes it. No generated stubs or proto files are needed.
  • Calls are unary.
  • A non-OK gRPC status is not an error: it lands in status_code / status_message for assertions.
  • Payloads are marshaled via protojson; responses use UseProtoNames + EmitUnpopulated, so assert with snake_case field names and expect zero-valued fields to be present.
  • A run opens one connection per target (address and TLS setting) and resolves each method by reflection once, then reuses both for every later call, including under load.

The step key is db:. sql: is the original spelling and still loads, so existing suites keep working; prefer db: in new tests.

Declare connections once at suite level and reference them by name, so a DSN (and its credentials) is written in one place:

databases:
erp:
driver: postgres
connection: "{{db_url}}" # templated at execute time like any config string
analytics:
driver: clickhouse
connection: "clickhouse://user:pass@localhost:9000/appdb"
tests:
- name: api-matches-db
steps:
- name: verify-row
db:
use: erp # supplies driver + connection
query: "SELECT id, name FROM users WHERE id = $1"
params: ["{{user_id}}"]
assertions:
- row_count == 1

databases: resolves up the suite tree exactly as data: does, so a parent suite.yaml can declare one connection every descendant reuses. Names resolve at load time: an unknown name fails the load and lists the ones that exist, rather than reaching the driver with an empty DSN.

use: and driver:/connection: are mutually exclusive: use: already supplies both.

By default every step opens a connection and closes it again. A connection that sets reuse: true is instead kept open for the whole run and shared by every step that names it:

databases:
erp:
driver: postgres
connection: "{{db_url}}"
reuse: true
max_open_conns: 4 # optional; zero leaves database/sql's default
max_idle_conns: 2
conn_max_lifetime: 5m

Reuse is opt-in on purpose: holding connections open changes what the server sees, so a suite that asserts on pg_stat_activity, or runs against a tight connection budget, keeps the open/close behavior by not asking for anything.

The cache is keyed by driver + rendered DSN, so one named connection whose DSN embeds a per-test variable still gets a handle per distinct database. Bounds apply when the handle is first created; redeclaring the same DSN with different bounds does not reconfigure a live pool. Handles are closed when the run finishes.

A reused sqlite :memory: connection persists across steps, which is often what you want; the default closes it between steps and each step starts empty.

Comparing a numeric column to a templated value needs the template unquoted: rows[0].id == {{user_id}} compares numbers, while rows[0].id == "{{user_id}}" compares a number to a string and fails.

steps:
- name: verify-row
db:
driver: postgres # postgres|postgresql|pg, mysql, clickhouse, duckdb, sqlite|sqlite3
connection: "postgres://user:pass@localhost:5432/db?sslmode=disable"
query: "SELECT id, name FROM users WHERE id = $1"
params: ["{{user_id}}"] # positional only
timeout: 5s # default 10s
response:
variables:
db_name: "rows[0].name"
total: "row_count"
assertions:
- row_count == 1
- rows[0].name == "{{username}}"
- name: seed
db:
use: erp
execute: "INSERT INTO users (id, name) VALUES ($1, $2)"
params: ["{{user_id}}", "{{username}}"]
assertions:
- rows_affected == 1
  • query (SELECT → rows, row_count, columns) and execute (DML → rows_affected) are mutually exclusive.
  • Empty connection opens an in-memory DuckDB: handy for pure-SQL smoke tests.
  • Parameters are positional: params is a list, in the driver’s placeholder order.
  • rows keeps at most a 1000-row sample; []byte columns coerce to string, timestamps to RFC3339Nano.
  • DB errors mark the step errored (with the error in evidence) rather than aborting the run.

Every other step is request/response. A stream is not: it stays open, messages arrive on their own schedule, and what you want to assert is “the event I care about turned up, in time, and looked right”.

A streaming step has two halves that read as a conversation:

  • triggers:: the things you do. Named, so a script can fire one.
  • expect:: the things the server says back, in order. Each one selects a message, validates it, and can run a script that fires the next trigger.
steps:
- name: order-lifecycle
sse:
url: "{{baseurl}}/events"
timeout: 30s
triggers:
- name: create # fires at connect
http:
method: POST
url: "{{baseurl}}/orders"
headers: {Content-Type: application/json}
body:
payload: '{"item":"widget","qty":2}'
save_as: created # the POST's own response
- name: cancel # idle until a script fires it
when: call
http:
method: POST
url: "{{baseurl}}/orders/{{order_id}}/cancel"
expect:
- match: "event._kind == 'order.created'" # SELECT the message
timeout: 5s
save_as: ev_created
assertions: # VALIDATE it: all of them run
- event.id == {{created.response.id}}
- event.qty == 2
- event.total > 0
- event._offset_ms < 500
script: # ANSWER BACK
language: js
execute: |
stream.trigger("cancel", {order_id: event.id});
- match: "event._kind == 'order.cancelled'"
timeout: 5s
assertions:
- event.id == {{ev_created.id}}
assertions:
- message_count == 2 # whole-stream facts

ws: is the same, plus the socket-only fields, send:, subprotocols:, on_close:, and message as the alias for event.

FieldMeaning
matchan assertion expression evaluated against each arriving message until one satisfies it. Empty matches any message
assertionsexpressions validating the message that matched. Every one runs
scriptruns after the assertions pass, against that message. Usually stream.trigger(...)
countwait for this many matches rather than one; save_as then captures the list, and assertions/script run for each
timeoutdeadline for this wait, measured from when it started waiting, not from the start of the step; falls back to the step’s timeout
save_ascapture the matching message into the test scope

match selects; assertions validate. The distinction decides what happens to a wrong message, and it is the most important thing on this page:

  • A message that fails match is someone else’s traffic: recorded and skipped, and the wait continues. Real subscriptions carry heartbeats and unrelated events, and failing on the first one would make them untestable.
  • A message that matched but fails assertions fails the step. The event the test asked for did arrive and it was wrong. Resuming the search would hunt for a second one that may never come and report a timeout: naming the wrong problem entirely:
expect #1 matched [1 @ 3ms] order.created but it failed validation:
event.total > 0 — expected > 0, got 0

A JSON object’s fields sit at the top level, so message.type and event.total read naturally. The envelope is always present under underscore-prefixed names, whatever the payload was:

FieldMeaning
_seq1-based arrival order
_offset_msmilliseconds since the connection opened
_kindtext/binary for websocket; the event name for SSE
_idthe SSE event id
_rawthe payload as received
datathe parsed payload, or the raw string when it is not JSON

_offset_ms is why arrival is recorded at all: “the event arrived” is rarely the question: “it arrived within 200ms of the one before” is.

message and event are the same thing under two names; use whichever suits the protocol you are thinking about. Both match: and assertions: are Liquid-rendered first, so they can refer to anything an earlier step, trigger or wait produced.

After the waits resolve, the step’s own assertions: see message_count (everything that arrived, including messages past the retention bound), messages, connected, status_code for SSE, and every name captured by a save_as. Use them for facts about the stream as a whole; use a wait’s assertions: for facts about one message.

  • A wait that is never satisfied fails the step.
  • A wait whose assertions fail fails the step.
  • A connection that could not be opened errors it: the question was never asked.

An unmet wait reports what it wanted and what arrived instead, because a bare “timeout” on a busy stream tells you nothing:

expect #1 timed out after 300ms waiting for "event._kind == 'order.created'" — saw instead:
[1 @ 5ms] heartbeat {"n":1}
[2 @ 11ms] heartbeat {"n":2}

A trigger is http: or script:, and fires either at connect or when a script calls it.

FieldMeaning
nameidentifies it for stream.trigger(name, params). Required for when: call
whenconnect (default) fires in order once the stream is open; call sits idle until a script fires it
httpa request: the same block an http: step takes, so auth, bodies, uploads and templating all work
scripta script instead, with the same task namespaces a script: step has. Mutually exclusive with http
save_ascaptures the outcome: for http, {status_code, response, headers, duration_ms}; for script, whatever it returned
repeatfire it this many times in sequence. save_as then captures the list
delaywait this long before firing. Rarely needed; reach for it when a server registers the subscription asynchronously after the handshake

trigger: (singular) is shorthand for a one-element triggers:. Setting both is a load error rather than a silent choice between them.

when: connect triggers fire in order, and each sees what the ones before it captured, so the second can address what the first created. A failure stops the sequence: amending an order that was never created just produces a second, more confusing error on top of the real one. The step reports trigger "amend" (#2): ... rather than the timeout it caused.

To fire the same trigger N times, use repeat:, not N copies of the block:

triggers:
- name: create
repeat: 3
http: {method: POST, url: "{{baseurl}}/orders"}
save_as: made # a list of 3
expect:
- match: "event._kind == 'order.created'"
count: 3 # order-independent, unlike three waits
assertions:
- event.total > 0 # runs against each of the three

Inside an expect: script, stream is in scope:

stream.trigger("cancel", {order_id: event.id});

The params become variables when that trigger’s config is templated, so the trigger above can write {{order_id}}. It returns {success: true, fired: "cancel"}, or {success: false, error: ...} naming the triggers that do exist if you mistyped the name: a silently ignored call would surface as an unexplained timeout much later.

It returns immediately. The request goes out on its own goroutine, so the reader is not stalled and the offsets of the messages it provokes still measure the server rather than the runner. There is no synchronous variant: the point of a stream step is that the response comes back as a message, which is what the next expect: is for.

The script itself runs synchronously on the reader, so keep it short: a slow script delays later messages and inflates their offsets. It can also test.set_variable() into the test scope, and test.fail() to fail the step.

Correlating against a trigger’s capture is safe even when the push wins the race. A service that publishes inside the same transaction that answers the request can deliver the event to a subscriber before the runner has finished reading the POST’s own response. A match: or an assertions: entry that reads a variable the triggers have not set holds its message until the triggers have settled, rather than templating {{created.response.id}} against nothing. A wait whose variables are all already resolved is judged the instant its message lands.

A non-2xx trigger does not fail the step: “POST returns 409 and the service pushes a conflict event” is a real test. But because an unexplained timeout is so often a trigger that did not do what the author assumed, the status is carried into the timeout message:

expect #1 timed out after 600ms waiting for "event._kind == 'order.created'" — no messages arrived while waiting
— trigger #1 POST http://localhost:8080/orders returned 422 Unprocessable Entity

A trigger that could not be sent at all is different, as is a script trigger that calls test.fail(): the step errors immediately rather than waiting out a deadline for a push that provably cannot come.

max_messages (ws) / max_events (sse) bound what is retained as evidence, default 1000; the count stays accurate beyond it. ws: also takes subprotocols:, send:, on_close: and tls:; sse: takes method:, query:, body: and tls: like an HTTP step.

Neither client reconnects. A silent reconnect would hide exactly the disconnection a test may exist to catch.

steps:
- name: custom-check
script:
language: js # default js. Also: js:v8, lua, starlark, cel, expr
execute: |
const r = response; // parsed body of the previous HTTP step
if (!r.items.every(i => i.price > 0)) {
test.fail("non-positive price found", 400);
}
test.set_variable("item_count", r.items.length);
test.success("all items valid", 200);
timeout: 10s # default 5s
params: { max: 10 } # extra vars merged over the scope

execute (inline) and file: (path relative to the test YAML’s directory) are mutually exclusive. See Scripting for the full test.* API, injected variables, and pass/fail semantics.

A playwright: step drives a browser with Playwright, under Node. It runs in the same test as HTTP, database and CLI steps, so one test can create data through the API, check it in the UI, and confirm it in the database.

Playwright comes from the project: the step runs Node in its working directory, and playwright or @playwright/test must resolve from there. The browser is Obscura unless you choose another: a lightweight headless browser that Playwright drives over the Chrome DevTools Protocol (CDP). Install Playwright in the project, and Obscura on the machine that runs the tests, with obscura and obscura-worker side by side on PATH (or TALOS_OBSCURA naming obscura):

Terminal window
npm install -D @playwright/test
curl -L https://github.com/h4ckf0r0day/obscura/releases/download/v0.2.3/obscura-x86_64-linux.tar.gz | tar xz -C ~/.local/bin

For a browser Playwright launches itself, install it with npx playwright install chromium (or firefox, webkit) and choose it as below.

Before a run starts, the runner checks that Node is on PATH, that Playwright resolves, and that Obscura is installed, for every browser step the run will reach, and stops with what to do when one is missing.

BrowserHow it runs
obscura (the default)The runner starts obscura serve for the step on 127.0.0.1, with loopback and private-network addresses allowed so it reaches the service under test, connects Playwright to it over CDP, and stops it when the step ends.
chromium, firefox, webkitPlaywright launches its own browser.
cdp: an endpointConnects to a browser that is already running and starts nothing: a shared Obscura server (obscura serve --port 9222 --allow-private-network serves ws://127.0.0.1:9222/devtools/browser), or a Chrome started with --remote-debugging-port.

The most specific choice wins:

  1. the step’s browser: or cdp:;
  2. the nearest suite’s playwright: block, which applies to every browser step in the suite and the suites below it;
  3. the run’s --browser (default $TALOS_BROWSER);
  4. obscura.
suite.yaml
playwright:
browser: chromium # or: cdp: ws://browsers.internal:9222

A spec step runs the browsers of the project’s playwright.config unless one of these chooses obscura or a cdp: endpoint. The runner then hands the endpoint to the specs as TALOS_CDP_ENDPOINT, and the specs use it through a fixture they import test and expect from:

fixtures.js
const base = require('@playwright/test');
exports.test = base.test.extend({
browser: [async ({ playwright }, use) => {
const endpoint = process.env.TALOS_CDP_ENDPOINT;
const browser = endpoint
? await playwright.chromium.connectOverCDP(endpoint)
: await playwright.chromium.launch();
await use(browser);
await browser.close();
}, { scope: 'worker' }],
});
exports.expect = base.expect;

On Obscura, page.setContent() in a script step waits for the content to be committed (waitUntil: 'commit'), because Obscura reports no load event for a page filled in place; in a spec, pass { waitUntil: 'commit' } or open a data: URL instead. The browser state an earlier step left carries into the next script step as cookies and local storage, and into a spec run as cookies.

A step takes one of two forms.

script: (or file:) is a JavaScript body run in a fresh browser context. await works at the top level.

steps:
- name: login-via-api
http:
method: POST
url: "{{app_url}}/api/login"
headers: { Content-Type: application/json }
body:
payload: '{"email": "{{email}}", "password": "{{password}}"}'
- name: dashboard-shows-the-user
playwright:
base_url: "{{app_url}}"
script: |
await page.goto("/dashboard");
await expect(page.getByTestId("user-name")).toHaveText(vars.username);
talos.set("order_count", await page.getByTestId("order").count());
talos.log("dashboard at", page.url());
assertions:
- passed == 1
- console_errors.length == 0
- name: orders-match
db:
use: app
query: "SELECT count(*) AS n FROM orders WHERE user_email = '{{email}}'"
assertions:
- rows[0].n == {{order_count}}
In scopeWhat it is
page, context, browserPlaywright’s page, browser context and browser
expect@playwright/test’s expect with auto-waiting matchers (present when @playwright/test is installed)
varsthe test’s variables
talos.set(name, value)set a variable for the steps after this one; export: publishes it to suite or run memory
talos.log(...)a line printed under the step and kept in JUnit system-out
requireNode’s require, resolving from the step’s working directory

A thrown error or a failing expect fails the step with its message. The step errors instead when it cannot run at all: Node missing, Playwright not installed, a browser that does not launch.

spec: runs existing Playwright Test specs, the same as npx playwright test, with the project’s playwright.config applying. The runner reads the result from Playwright’s JSON report.

steps:
- name: checkout-journey
playwright:
spec: e2e/checkout.spec.ts # file, directory or pattern; `.` runs every spec
grep: "@smoke"
project: chromium
base_url: "{{app_url}}"
assertions:
- failed == 0
- flaky == 0

A spec reaches the test’s variables and hands values back through environment variables:

VariableWhat it is
TALOS_VARSpath of a JSON file holding the test’s variables
TALOS_EXPORTpath to write a JSON object to; its keys become variables for the steps after this one
TALOS_BASE_URLthe step’s base_url
TALOS_STORAGE_STATEpath of a storage state holding the cookies the test’s HTTP steps collected for base_url and the session an earlier browser step left; set when there is one
playwright.config.ts
use: {
baseURL: process.env.TALOS_BASE_URL,
storageState: process.env.TALOS_STORAGE_STATE,
},
KeyFormDefaultMeaning
script / filescript—The body, inline or from a file relative to the test file.
specspec—Spec files to run, as playwright test takes them.
grep / project / configspec—--grep, --project, --config.
browserbothobscuraobscura, chromium, firefox or webkit; see Choosing the browser. A spec step takes only obscura.
cdpboth—The DevTools endpoint of a browser that is already running, to connect to instead.
headlessbothtruefalse shows the browser (--headed in spec form).
base_urlboth—Resolves relative page.goto() paths, and is the origin whose cookies are carried in from HTTP steps.
viewportscriptPlaywright’s{ width, height }.
tracebothon_failurealways, on_failure or never. In spec form, --trace on, retain-on-failure or off.
screenshotscripton_failureA full-page screenshot when the script ends.
videoscriptneverA video of the session.
cwdboththe test file’s directoryWhere Node runs and node_modules resolves from.
envboth—Extra environment variables, templated.
nodebothnodeThe Node.js binary.
timeoutboth5mCaps the whole run.
exportboth—Publish the variables the step set to suite or run memory.

Keys that apply to the other form are load errors: in spec form, the viewport, screenshots, video and any browser other than Obscura come from playwright.config.

The steps of a test share one session:

  • The HTTP steps’ cookies for base_url are added to the browser, so an API login carries into the UI.
  • The browser’s cookies go back into the test’s cookie jar afterwards, so an HTTP step after a UI login is logged in too.
  • A script step’s browser state (cookies and local storage) carries to the test’s next browser step, so each step starts where the last one left off. A retried test starts a fresh session.
FieldWhat it is
statuspassed or failed
passed / failed / skipped / flakytest counts (script form: one test)
testseach test: title, file, project, status, duration_ms, retries, error
errorsfailure messages, and errors Playwright reported outside a test
console_errors / page_errorswhat the page logged with console.error, and uncaught exceptions (script form)
url / titlethe page’s when the script finished (script form)

Variables the step set are in scope too.

The report, traces, screenshots and videos are kept in the run’s artifacts directory (see Outputs). Open a trace with npx playwright show-trace <path>. The step’s result is stored in step_evidence with --db.

kis test load runs a suite’s other step types. A browser per virtual user would measure the load generator rather than the service, so a load profile that reaches a browser step stops before it starts and names the step; exclude browser tests from load runs with --exclude-tags.

Hooks are steps that run at suite boundaries; they use the identical step schema.

HookRunsFailure behavior
before_allonce per suite, before its teststhe suite’s tests do not run; each is reported errored, naming the failed step. Child suites and after_all still run
before_eachfull suite chain root → leaf, before every test (and every table row)test marked errored, steps skipped; after_each still runs
after_eachleaf → root, after every test, pass or faillogged, does not change the test result
after_allonce per suite, always: even after errorslogged, does not change the suite result

Variables set by before_all persist in the suite scope (visible to all tests in the suite and to child suites). Variables set by before_each live in the test scope. This is how auth tokens propagate:

suite.yaml
before_each:
- name: authenticate
http:
url: "{{baseurl}}/auth"
method: POST
body:
payload: '{"client_id": "{{client_id}}"}'
response:
variables:
token: "response.access_token"

Suite-level hooks (before_all/after_all) get their own cookie jar; per-test hooks share the test’s jar.

Hooks are declared in suite.yaml, or at the top of a file run on its own with -t file.yaml. A hook in any other file of a directory suite fails the load; see Where suite-level keys live.

A step with steps: instead of a protocol is a group: its children run in order, in the test’s scope, so a variable one child extracts is visible to the next and to the steps after the group.

steps:
- name: login
retries: 2
timeout: 20s
steps:
- name: get-token
http:
method: POST
url: /auth/token
response:
variables:
token: "response.access_token"
- name: check-token
http:
url: /auth/verify
headers: { Authorization: "Bearer {{token}}" }
assertions:
- status_code == 200
assertions:
- token != null
- name: create-order
http:
method: POST
url: /orders
headers: { Authorization: "Bearer {{token}}" }
  • The group passes when every child passes. The first failing child ends the group unless the child sets continue_on_error; the group then fails, and its error names the child.
  • timeout caps the whole group, retries runs the whole group again when it errors, and continue_on_error on the group lets the test go on past it.
  • assertions: on a group run after its children pass, against the scope they left.
  • Children can be groups. The console prints a group’s children indented beneath it, the group’s line says how many passed (2/2 steps passed), and --db stores each child in step_runs with the group as its parent_step_id.

A step block used by many tests lives in a shared file and is pulled in with import:. Keep it in a _-prefixed directory so the loader does not treat it as a suite.

tests/_shared/auth.yaml
steps:
- name: login
steps:
- name: get-token
http: { method: POST, url: /auth/token, response: { variables: { token: "response.access_token" } } }
- name: check-token
http: { url: /auth/verify, headers: { Authorization: "Bearer {{token}}" } }
- name: health
http: { url: /health }
tests/orders/checkout.yaml
import:
shared: ../_shared/auth.yaml
tests:
- name: checkout
steps:
- insert: $shared:/steps[name=login] # the login group, as one step
- splice: $shared:/steps # or every step in the file, in order
- name: pay
http: { method: POST, url: /pay }

insert: and splice: both place the selected steps into the list; a group arrives as one step with its children. Paths are relative to the importing file. Hooks take shared steps the same way.

Terminal window
kis test run -t tests/ --tags smoke # tagged smoke
kis test run -t tests/ --tags "smoke,api" # smoke AND api
kis test run -t tests/ --tags smoke --tags nightly # smoke OR nightly
kis test run -t tests/ --tags api --exclude-tags wip

Commas inside one --tags value = AND; repeating the flag = OR; --exclude-tags uses the same syntax and wins ties. A test’s effective tags are the union of its own tags and every ancestor suite’s tags.

tests:
- name: refund-to-card
skip: "the sandbox has no card processor (PAY-12)"
steps: [...]
- name: new-checkout-flow
only: true # while writing it: run just this one
steps: [...]
  • skip: true or skip: "<reason>" on a test, or in a suite.yaml for the whole suite, reports the test as skipped with the reason. The console shows it, and JUnit carries it as <skipped message="...">.
  • When any test in the run sets only: true, only those tests run; the console says how many. --forbid-only makes such a run fail instead, which is the setting for CI.
  • A suite’s hooks run only when at least one test in the suite or below it will run. Narrowing a run with --tags, only: or skip: does not set up suites none of whose tests run.
  • kis test load leaves skipped tests out of the workload and honours only: the same way.

--dry-run loads the suite tree, validates it, and renders every step’s templates without running any step. Each test reports as skipped; a step whose template does not render reports as errored.

What a step would set when it runs stands in as a placeholder, <name>, for the steps after it: its response.variables, a stream’s save_as, the names a script passes to set_variable, set_suite_var or set_run_var, and the names a browser script passes to talos.set, published wherever the step’s export: or setter would put them. A dry run therefore shows each request as it would be sent, with <order_id> where the extracted id goes, and an undefined variable it reports is one nothing sets: a misspelling, or a value the environment does not provide. A name a script computes at run time cannot be known in advance.

A step declares exactly one protocol. A step that declares two (http: and cli:, say) is a load error naming both.

Loading is strict: a key the schema does not know fails the load.

This matters more than it sounds. The classic mistake is indenting assertions: inside the http: block instead of beside it. Under a lenient parser the key is silently dropped, the step has nothing to assert, and every run reports a pass, including with a deliberately wrong status_code. A test that quietly asserts nothing is worse than one that fails.

# Wrong: assertions belongs to the step, not the http config
- name: ready-check
http:
url: "{{baseurl}}/ready"
assertions:
- status_code == 200
line 9: field assertions not found in type types.HTTPConfig
"assertions" belongs at step level — unindent it to line up with "http:", not inside it

Where a misplaced key is valid at an outer level the error says so; otherwise it lists the keys that are valid where you wrote it. The same strictness catches export: targets other than suite/run, a use: naming a connection that is not declared, a step that sets both db: and sql:, and a suite-scoped key (variables:, data:, databases:, a hook) written in a test file of a directory suite, where it would otherwise never apply.

Everything else load-time validation enforces: every test has a name, every step has a name and exactly one protocol key or steps:, hooks the same way, and every database step resolves to a driver plus a connection. On a cli: step it also enforces command: XOR shell:, and rejects args: beside shell:: sh -c ignores anything appended, so those arguments would simply never be passed.