Skip to content
Talk to our solutions team

Data Observability

data.svc emits four things: metric instruments on the engine and on every route, one span per request, structured logs on stdout, and a set of admin routes that report live runtime state. Only the last two reach you without extra wiring.

SignalEmittedLeaves the process
Engine metrics: operations, transactions, hook errorson every operationno exporter is initialised
HTTP server metrics: count, in-flight, bytes, latencyper wrapped routeno exporter is initialised
Request spansper wrapped routeno exporter is initialised
Logsalwaysyes: stdout
Runtime state (queues, locks, tenant inventory)on requestyes: admin routes
Compiled SQL for one requeston requestyes: ?stats=true
Engine debug logger (AST + SQL + timings)library onlynot available in this service
EXPLAIN / query planneverdoes not exist

The instruments below are registered whether or not an exporter exists. Without one, meter and tracer lookups fall back to the global no-op providers and every measurement is discarded, at the cost of its atomic add. Naming a collector builds the real providers at boot, before the first engine, so the instruments record into them:

telemetry:
collector: otel-collector:4317 # required; unset keeps every signal no-op
environment: production
insecure: false
traces:
enabled: true
sampling_rate: 0.1
KeyDefaultEffect
telemetry.collector—OTLP gRPC endpoint. Unset is the only thing that keeps the signals no-op, and the service says so once at boot
telemetry.environmentproductionResource attribute on every signal
telemetry.insecurefalsePlain gRPC to the collector instead of mTLS
telemetry.traces.enabledfalseSpans. Metrics and logs export without it
telemetry.traces.sampling_rate0.1Fraction of traces sampled
telemetry.local_devfalseRelaxed local-development exporter settings

telemetry.enabled is not a key and is refused at boot. A collector that cannot be reached is logged and does not fail the boot: a service that cannot export telemetry still serves data. On shutdown the providers are flushed, so a draining process exports its last measurements.

The engine reports lifecycle events to one process-wide observer, built once and shared by every per-tenant engine. Ten instruments are registered. If the meter is unavailable or any registration fails, the observer degrades to a no-op for the life of the process and logs the reason at error level, engine construction never fails on telemetry.

MeasuresKindUnitLabelsFires
Engine operationscounter—operation, entity, statusend of every query, create, update, delete, restore
Operation latencyhistogrammsoperation, entity, statussame call; covers hooks, compile, driver and result scan
Rows returned or affectedhistogram—operation, entitysame call, skipped when the count is unknown
Transactionscounter—committed, statusend of every transaction
Transaction latencyhistogrammscommitted, statusbegin through commit or rollback
Hook errorscounter—hook, phase, operation, entityonly when a hook returns an error
Migration runscounter—nonelibrary migration path only
Schema objects createdcounter—nonelibrary migration path only
Migration step failurescounter—nonelibrary migration path only
Migration latencyhistogrammsnonelibrary migration path only

On the operation, latency and rows instruments, operation is one of query, create, update, delete, restore, and status is ok or error. phase is one of BeforeValidate, AfterValidate, BeforeExec, AfterExec, AfterScan. Successful hooks emit nothing, so hook-error count is a rate you can alert on directly.

Nested transactions report once, at the outer call, a commit failure and a rollback both arrive with committed=false; the status label separates a clean rollback from a driver error.

The four migration instruments never fire in data.svc. They are emitted by the library’s own create-if-missing migration call, which this service does not use, its plan and apply pipeline is a separate path with no observer attached.

Every route registered by the service is wrapped in a per-route instrumentation middleware. Five instruments are created on first use:

InstrumentKindRegistered unit
http.server.request_countcounter—
http.server.active_requestsup-down counter—
http.server.request_content_lengthcounterBy
http.server.response_content_lengthcounterBy
http.server.durationhistogramMs

Each carries http.method, http.scheme, user_agent.original and http.route (the route pattern, not the resolved path). When the request resolved a tenant, it also carries customer, product, env and tenant.

http.status_code is added to the attribute set only after the handler returns, so it is present on the two content-length counters and on http.server.duration, and absent from http.server.request_count.

/data/health and /data/ready are registered directly, ahead of the instrumentation wrapper, so they produce neither metrics nor spans. See Checking a tenant is servable for what the probes actually assert.

The wrapper carries a fallback that degrades to a plain passthrough, no metrics, no span, when the process resolves no service instance id. It never triggers in a booted service: the chassis registers a service-discovery client unconditionally, and when servicediscovery_url is unset that client is a stub whose service id is a non-empty placeholder. Every route is instrumented whether or not you set sid or servicediscovery_url; the placeholder id is what lands on the span as the server name.

One server span per instrumented request. The span name and the http.route attribute are both the route pattern. W3C traceparent and baggage headers on the incoming request are extracted, so the span joins a caller’s trace; a valid remote span already on the context is attached as a link. Tenant identity rides on the span as the same four attributes the metrics use, and any tags request headers are copied onto it. When the handler returns, the span records http.status_code and a span status derived from it. A second span is opened once at boot, around service start.

A panic inside a handler is recorded through the tracer by the router’s panic handler.

The engine adds nothing. It reports through metric events, not spans, so a captured trace is one span per request with no child spans for compile, hooks or the database round trip. Per-operation timing comes from the operation-latency histogram, or from ?stats=true on a single request.

Error responses carry no request id and no trace id. The response envelope is error, code and details only. Correlate a client-side failure with the logs by tenant, route and timestamp.

Logs are structured records written to stdout through a console writer, timestamped RFC 3339. Colour is applied only when stdout is a terminal, so a redirected stream or a supervisor-captured stream is clean. Every line carries service_id, datacenter and cluster, taken from the sid, datacenter and cluster configuration keys.

KeyTypeDefaultEffect
log.levelstringerrorMinimum severity: trace, debug, info, warn, error, fatal, panic

These conditions are logged and are invisible everywhere else. No probe fails, no request errors:

ConditionConsequence
A datastore init script failsLogged; the engine still serves traffic
The materialized-refresh runner fails to installLogged; mutations succeed and refresh tasks are never enqueued
Two schema documents declare the same entity or traitLogged at error; the first declaration stands and the later one is ignored
schemasource: is unset, unknown, or points at a missing directoryLogged as a warning; falls back to the metadata service
A tenant listed under loadtenants: fails to warmLogged; boot continues
The rate limiter cannot build, or a route names an undeclared profileLogged. Requests are served unlimited, so alert on this line
Engine metric registration failsLogged at error; engine metrics become no-ops
jwks.disabled: trueLogged once at error level, as a SECURITY: line. The service is accepting tokens without signature verification

Two logging knobs do nothing in this service, and both appear in shipped example configurations:

  • log.loglinenumbers is read by no code. The service builds its logger with source line numbers off, unconditionally.
  • realtimedebug, with the _debug_log_ query parameter or the X-Debug-Log header, switches the request context to a debug logger that this service never registers. Every log call in data.svc resolves the service logger directly, so the flag changes no output. Raise log.level instead.

Add ?stats=true to a REST read or write and the response envelope gains a stats block containing the compiled sql, one entry per relation eager-load subquery in include_sql[], duration_ms, and rows_affected on mutations. This is the only per-request SQL visibility that works against a deployed service, and it is available to any caller who can make the request.

Terminal window
curl -s 'https://api.example.com/data/rest/invoice?status=open&stats=true' -H 'Authorization: Bearer <token>' -H 'X-Customer: acme' -H 'X-Product: erp' -H 'X-Env: prod' -H 'X-Tenant: eu1'

The full flag set and the exact envelope are in Request flags.

What data.svc exposes is stats.sql, the compiled statement, the statement count, the row count and the elapsed milliseconds. Bound argument values are deliberately not included: they carry plaintext user input, and a query log that reproduces them becomes a second copy of your data with none of its access controls.

There is no EXPLAIN. No route, no query DSL keyword and no request flag returns a database query plan or a cost estimate, and the service issues no EXPLAIN of its own. To read a plan, take the statement from stats.sql and run it against the database yourself.

These routes report what the process is currently doing for the calling tenant. All require the four CEPT headers and a token.

RouteReportsGate
GET /data/admin/data/materialize/statsCross-entity refresh queue depth by status, plus cumulative runner counters. The runner block is absent when no runner exists for this tenantadmin
GET /data/admin/data/materialize/pendingPending and failed refresh tasks with attempt counts: page size fixed at 100admin
GET /data/admin/data/locksCoordination locks with holder, operation kind, acquisition and expiry, is_expired and seconds_to_expiry. Accepts prefix, op_kind and include_expiredadmin
GET /data/superadmin/tenantThe tenant keys this process currently holds live resources for. The closest thing to a cache inventorysuperadmin
GET /data/admin/data/hashesReturns an empty list; the catalogue is not enumerableadmin
GET /data/admin/data/cache/statsThe entity cache’s counters per declared backend: hits, misses, stores, invalidations, errors since the process startedadmin

Each /data/admin/data/… route above has a /data/superadmin/tenant/{tenant}/… twin, the same handler, re-rooted, reading the state of a named sibling tenant. The twin path drops the admin/data segment: GET /data/superadmin/tenant/{tenant}/locks. GET /data/superadmin/tenant itself has no admin equivalent, because an admin cannot list tenants it does not belong to. Migration plans, applied-step history and seed state are covered in Migrations.

Not availableDetail
Query plansNo EXPLAIN emitter exists anywhere in the service
Connection pool statisticsThe pool registry tracks reference counts and endpoint snapshots internally; no route exposes them
Per-tenant engine metricsTenant is not a label on any engine instrument, by design
The compiled-query cache catalogueThe route returns an empty list

Every statement the engine runs can be written to the service log, all of them, only the slow ones, or none, and the setting changes while the service runs. The boot block sets the starting point:

sqllog:
mode: slow # off (default) | slow | all
slower_than: 250ms # slow: record a successful statement at or above this
sample: 1 # all: the fraction of successful statements recorded, 0 to 1
max_length: 4096 # bytes of statement text per line

A line is one statement: the tenant and user it ran for, the operation and entity, the row count, the microseconds it took, the number of bound arguments, the statement text on one line, and the error when it failed. Bound values are never available to it. The engine hands the log a count, not the values, so a password hash or a token secret cannot end up in a log line through this switch. Failures are always recorded in slow and all; slow keeps successful statements above the bar and all keeps the sampled fraction. off records nothing but keeps counting.

The admin plane reads and changes the setting without a restart:

Terminal window
curl -H "Authorization: Bearer $ADMIN" -H "X-Customer: acme" -H "X-Product: shop" -H "X-Env: prod" -H "X-Tenant: main" \
http://data.example/admin/data/sqllog
# {"mode":"slow","slower_than":"250ms","sample":1,"max_length":4096,"seen":48211,"logged":37}
curl -X PUT -H "Authorization: Bearer $ADMIN" -H "Content-Type: application/json" ... \
-d '{"mode":"all","sample":0.1}' http://data.example/admin/data/sqllog

seen counts every statement since the process started and logged the ones written, so the gap between them is what the mode is filtering. A setting the log refuses. A mode it does not know, a sample outside 0 to 1, answers 400 and changes nothing. The setting is per process and applies to every tenant it serves; the tenant is a field on every line, not a switch. Recording every statement of a busy service is a lot of log: reach for all with a sample, or for a minute at a time, and leave slow on the rest of the time.

The security audit stream is durable and is read from the database. Every tenant engine writes a tamper-evident hash chain into its own data_audit_events table: one row per access denial (the access gate’s and the safety guard’s), per decrypt operation, per internal error, and, for the entities that opt in, per successful mutation or read, with the principal, operation, entity, outcome and reason lifted into columns and the exact event kept alongside its position in the chain (seq), its prev_hash and its hash. The chain’s head is the table itself: a row is appended under a per-tenant lock at the previous position plus one, so a restart or a second process continues the chain rather than starting another, a mutation that runs in a transaction, every operation of a bulk envelope, a transactional action, writes its audit row on that transaction, as its last statement: a rolled-back envelope leaves no audit rows, a committed one leaves exactly one per operation, a mutation outside a transaction writes its row after its own commit, and a denial is recorded even when the transaction it happened in rolls back. Nothing serves the stream over HTTP yet. Read it with SQL, and check it by walking the rows in seq order: positions run 1..n with no gap, each prev_hash is the previous row’s hash, and each hash is sha256 over the previous hash and the event.

Successful operations are recorded only for entities that opt in with an audit block, declared by the product. Denials are recorded whatever the block says; without a block nothing else is:

entities:
- name: payroll
audit:
mutations: true # off unless set; one row per successful create, update and delete
reads: true # off unless set; one row per read that names the entity
fields: names # names (default) records field names and row ids, never values; none records neither

Auditing is opt-in because every audited operation passes through the tenant’s chain in order: each append takes the chain lock, reads the head, inserts and commits before the next may start, so a tenant’s audited operations share one serial section of under a millisecond. Opting the benchmark’s write entity in cost 44% of its throughput and 20 ms on every write at p95; opting its read entity in halved throughput again. Opt in the entities whose access you must be able to reconstruct, keep reads: true for entities read rarely, and keep a tenant’s audited write rate well under a thousand a second. A read row carries the projected field names, or the returned columns for a default projection, the first hundred primary keys with a rows count in the event’s metadata, and the action name when the read ran inside one; every query that names the entity counts. A REST get, a DSL or GraphQL query, an include, a read inside an action or a script, a tenant overlay cannot change the block: what the stream records is not a tenant’s to decide. Per-entity history and change data capture remains the place for row-level before-and-after values; the audit stream records that an operation happened and who did it, never the values.

With telemetry configured, every engine operation is a span named seshat.<operation> under the request’s span, with entity and rows attributes, the real start and end times, and the error recorded when one occurred. seshat_access_denied_total counts the operations an access gate refused, by entity, operation and tier.