Skip to content
Talk to our solutions team

Rate limiting

The gateway limits inbound traffic per tenant, per client, or per anything else you can key a bucket on. It is configured in three layers, split by who owns the value: the deployment owns the engine, the routing model owns the binding, and a tenant’s own config owns its ceiling.

Traefik’s own rateLimit middleware stays reachable through passthrough, unchanged. Reach for this one when the limit should follow a tenant rather than a connection.

ratelimit: is a boot-plane key in the bootstrap file, alongside timeouts:. It carries a store DSN, so it is infrastructure, and like the other boot keys it is stripped from routing-directory files.

ratelimit:
enabled: true
on_store_failure: allow # allow (default) | deny
store:
kind: tiered # memory (default) | distrib | tiered
distrib:
backend: dragonfly
dsn: ${RATELIMIT_DSN} # credentials from the environment, never from YAML
tiered:
local_capacity_ratio: 0.5
sync_interval: 200ms
profiles:
api_default: { algorithm: token_bucket, capacity: 200, refill_rate: 200 }
auth_burst: { algorithm: fixed_window, window: 1m, limit: 30 }
ai_inflight: { algorithm: concurrency, max: 8 }
anon_by_ip: { algorithm: sliding_window, window: 1m, limit: 60 }

Algorithms are token_bucket, leaky_bucket, fixed_window, sliding_window and concurrency.

kindWhere the counters liveUse
memoryIn the processOne replica, or a limit that only has to be approximate
distribThe shared storeOne ceiling across every replica
tieredA local bucket reconciled against the shared storeProduction. The common path stays local, so a store blip costs accuracy rather than traffic

A kisai-ratelimit entry in an endpoint’s middlewares: list, like kisai-auth and kisai-timeout:

endpoints:
- name: to-account
path: { prefix: /account/ }
backend: account
middlewares:
- customer-domain-mapper: { default: true }
- kisai-ratelimit: { profile: auth_burst, key: ip }
- name: to-ai
path: { prefix: /ai/ }
backend: ai
middlewares:
- customer-domain-mapper: { default: true }
- kisai-ratelimit: { profile: ai_inflight }
FieldDefaultMeaning
profilerequiredA profile declared under ratelimit.profiles
keytenantWhat the bucket counts. See below
cost1How much one request spends
sourcecriterion—A raw source criterion, overriding key

A binding to a profile that is not declared fails the boot, and so does a binding made while rate limiting is disabled. A rate limit that silently is not there is a posture downgrade nobody notices until the incident.

keyBucket keyUse
tenantThe full CEPT tuple, customer:env:product:tenantAuthenticated, tenant-scoped traffic
ipThe client IPPre-auth and anonymous traffic
hostThe request HostPer-site limits
domainThe resolved routing domainBefore the domain map has resolved
header:<name>That header’s valueAn API-key id, a partner id

An unknown value, or header: with no name after the colon, is refused at boot with a message naming the endpoint.

Two consequences of the gateway being the edge:

key: tenant on a request with no CEPT headers passes through. The limiter does not invent a tenant. Protect pre-auth paths with key: ip.

key: ip reads the direct peer, never X-Forwarded-For. A bucket the caller can choose by rotating a header is not a rate limit. Honouring forwarded addresses is opt-in, by declaring sourcecriterion.ipstrategy with a depth or excludedips instead of using the key shorthand. That is a separate setting from an entrypoint’s trustedips, which governs which peers may set X-Forwarded-* at all.

A tenant’s own configuration may raise or lower a profile, and it takes effect without a restart:

ratelimit:
profiles:
api_default: { algorithm: token_bucket, capacity: 2000, refill_rate: 2000 }

The override registers against that tenant and resolves ahead of the fleet default. The bucket key keeps the base profile name, so raising a ceiling changes the limit rather than handing that tenant a fresh empty bucket.

on_store_failure decides what a request gets when the limiter cannot measure it.

ValueBehaviour
allow (default)The request is served, and the failure is logged loudly
denyThe request is refused with 503 and {"error":"rate_limit_unavailable"}

The default is allow because this is the public edge, where a store blip taking every route to 100% refusals is a worse outage than a window of unmetered traffic. Set deny on a limit that is a hard safety bound rather than a fairness one.

The refusal is 503 rather than 429 on purpose: the caller exceeded nothing, and a 429 would send them chasing a quota that is not the problem.

Over the ceiling, the route answers 429 in the same shape every kis.ai service uses, so a client cannot tell the edge from a service behind it.