Rate limiting
The gateway limits inbound traffic per tenant, per client, or per anything else you can key a bucket on. It is configured in three layers, split by who owns the value: the deployment owns the engine, the routing model owns the binding, and a tenant’s own config owns its ceiling.
Traefik’s own rateLimit middleware stays reachable through
passthrough, unchanged. Reach for this one when the limit
should follow a tenant rather than a connection.
1. The engine
Section titled “1. The engine”ratelimit: is a boot-plane key in the bootstrap file, alongside timeouts:. It carries a
store DSN, so it is infrastructure, and like the other boot keys it is stripped from routing-directory
files.
ratelimit: enabled: true on_store_failure: allow # allow (default) | deny store: kind: tiered # memory (default) | distrib | tiered distrib: backend: dragonfly dsn: ${RATELIMIT_DSN} # credentials from the environment, never from YAML tiered: local_capacity_ratio: 0.5 sync_interval: 200ms profiles: api_default: { algorithm: token_bucket, capacity: 200, refill_rate: 200 } auth_burst: { algorithm: fixed_window, window: 1m, limit: 30 } ai_inflight: { algorithm: concurrency, max: 8 } anon_by_ip: { algorithm: sliding_window, window: 1m, limit: 60 }Algorithms are token_bucket, leaky_bucket, fixed_window, sliding_window and concurrency.
Choosing a store
Section titled “Choosing a store”kind | Where the counters live | Use |
|---|---|---|
memory | In the process | One replica, or a limit that only has to be approximate |
distrib | The shared store | One ceiling across every replica |
tiered | A local bucket reconciled against the shared store | Production. The common path stays local, so a store blip costs accuracy rather than traffic |
2. Binding a profile to an endpoint
Section titled “2. Binding a profile to an endpoint”A kisai-ratelimit entry in an endpoint’s middlewares: list, like kisai-auth and
kisai-timeout:
endpoints: - name: to-account path: { prefix: /account/ } backend: account middlewares: - customer-domain-mapper: { default: true } - kisai-ratelimit: { profile: auth_burst, key: ip }
- name: to-ai path: { prefix: /ai/ } backend: ai middlewares: - customer-domain-mapper: { default: true } - kisai-ratelimit: { profile: ai_inflight }| Field | Default | Meaning |
|---|---|---|
profile | required | A profile declared under ratelimit.profiles |
key | tenant | What the bucket counts. See below |
cost | 1 | How much one request spends |
sourcecriterion | — | A raw source criterion, overriding key |
A binding to a profile that is not declared fails the boot, and so does a binding made while rate limiting is disabled. A rate limit that silently is not there is a posture downgrade nobody notices until the incident.
What a bucket counts
Section titled “What a bucket counts”key | Bucket key | Use |
|---|---|---|
tenant | The full CEPT tuple, customer:env:product:tenant | Authenticated, tenant-scoped traffic |
ip | The client IP | Pre-auth and anonymous traffic |
host | The request Host | Per-site limits |
domain | The resolved routing domain | Before the domain map has resolved |
header:<name> | That header’s value | An API-key id, a partner id |
An unknown value, or header: with no name after the colon, is refused at boot with a message
naming the endpoint.
Two consequences of the gateway being the edge:
key: tenant on a request with no CEPT headers passes through. The limiter does not invent a
tenant. Protect pre-auth paths with key: ip.
key: ip reads the direct peer, never X-Forwarded-For. A bucket the caller can choose by
rotating a header is not a rate limit. Honouring forwarded addresses is opt-in, by declaring
sourcecriterion.ipstrategy with a depth or excludedips instead of using the key shorthand.
That is a separate setting from an entrypoint’s trustedips,
which governs which peers may set X-Forwarded-* at all.
3. Per-tenant ceilings
Section titled “3. Per-tenant ceilings”A tenant’s own configuration may raise or lower a profile, and it takes effect without a restart:
ratelimit: profiles: api_default: { algorithm: token_bucket, capacity: 2000, refill_rate: 2000 }The override registers against that tenant and resolves ahead of the fleet default. The bucket key keeps the base profile name, so raising a ceiling changes the limit rather than handing that tenant a fresh empty bucket.
When the store cannot answer
Section titled “When the store cannot answer”on_store_failure decides what a request gets when the limiter cannot measure it.
| Value | Behaviour |
|---|---|
allow (default) | The request is served, and the failure is logged loudly |
deny | The request is refused with 503 and {"error":"rate_limit_unavailable"} |
The default is allow because this is the public edge, where a store blip taking every route to
100% refusals is a worse outage than a window of unmetered traffic. Set deny on a limit that is a
hard safety bound rather than a fairness one.
The refusal is 503 rather than 429 on purpose: the caller exceeded nothing, and a 429 would
send them chasing a quota that is not the problem.
What a limited caller sees
Section titled “What a limited caller sees”Over the ceiling, the route answers 429 in the same shape every kis.ai service uses, so a
client cannot tell the edge from a service behind it.
See also
Section titled “See also”- Configuration: every boot key
- Passthrough: Traefik’s own
rateLimitmiddleware - Security: the rest of the edge posture