Skip to content
Talk to our solutions team

Load balancing

A backend with more than one URL round-robins by default. Anything beyond that is a named profile, declared once and opted into per backend. The underlying primitives are Traefik’s.

Terminal window
gateway.svc config generate load-balancing
loadbalance:
default: # applied to any backend with no explicit profile
algo: wrr
healthcheck: { path: /health, interval: 30s, timeout: 5s }
profiles:
sticky-session:
algo: sticky
sticky: { cookie: srv-id, httpOnly: true, secure: true }
backends:
- backend:
name: api
urls: ["https://10.0.0.11:7102", "https://10.0.0.12:7102"]
profile: sticky-session
algoShapeUse
wrrapplies to the backend’s own URLsWeighted round-robin. The default
stickyapplies to the backend’s own URLsRound-robin plus a session cookie
weightedwraps other backendsA fixed traffic split
mirrorwraps other backendsOne backend serves; others get a sampled copy
failoverwraps other backendsA primary with a health-gated standby
adaptivewraps other backendsWeights recomputed from reported load

The last four are compound: they build a service over other named backends rather than over a list of URLs. A compound profile’s members must be backends that are themselves declared, and a cycle through profiles and backends is refused at load.

sticky-session:
algo: sticky
sticky:
cookie: srv-id
httpOnly: true
secure: true
sameSite: lax

Set secure: true whenever the route is served over TLS, and prefer httpOnly: true so page scripts cannot read the affinity cookie.

weighted-80-20:
algo: weighted
members:
- { backend: prod, weight: 4 }
- { backend: canary, weight: 1 }

Weights are relative, not percentages. A negative weight is refused.

shadow-10pct:
algo: mirror
primary: prod
mirrors:
- { backend: canary, percent: 10 }

The primary serves the request and its response is what the client gets. Mirrored copies are fire-and-forget and their responses are discarded, which makes this the right shape for exercising a new build against real traffic without putting it in the response path.

region-failover:
algo: failover
primary: us-east
fallback: us-west
healthcheck: { path: /health/deep, interval: 10s, timeout: 3s }

Traffic goes to the primary while its health check passes, and to the fallback when it does not. The health check is what drives the flip, so a failover profile without one has nothing to act on.

Session affinity and failover pull against each other: a flip moves a session to a backend that has never seen it. Decide which one the workload actually needs.

An adaptive profile polls each member’s own metrics endpoint and rewrites the weights between them. From the router’s point of view it is an ordinary weighted service whose weights happen to change.

adaptive-pool:
algo: adaptive
members: [{ backend: pool-a }, { backend: pool-b }, { backend: pool-c }]
reconcile:
interval: 10s
metrics_path: /health/load
timeout: 3s
formula: cpu_then_memory
min_weight: 1
max_weight: 100
scale: 100
KeyDefaultMeaning
membersrequiredThe backends to balance between
reconcilerequiredThe polling loop
reconcile.interval10sPoll period. Values below 5s are refused
reconcile.metrics_path/health/loadWhere each member reports its load
reconcile.timeout, .scheme, .port, .headersHow the probe is made
formulacpu_then_memoryHow a weight is derived
min_weight1Floor. Must be at least 1
max_weight100Ceiling. Must be at least min_weight
scale100The multiplier a formula works against

Each member answers metrics_path with JSON. Every field is optional, and a formula reads only what it needs:

{
"cpu_usage_percent": 42.0,
"memory_usage_percent": 61.5,
"available_cores": 2.5,
"available_memory_mb": 3072,
"active_requests": 17
}
formulaWeight before clamping
cpu_then_memoryscale × (1 - max(cpu_usage_percent, memory_usage_percent) / 100)
inflight_countscale / (1 + active_requests)
available_capacityscale × (available_cores × 10 + available_memory_mb / 1024)

The result is rounded and clamped into [min_weight, max_weight]. A formula name outside this set is refused at load, with the accepted names in the message.

cpu_then_memory sheds load from whichever resource is tighter. inflight_count is the one to reach for when request cost is uneven and queue depth is the better signal. available_capacity suits a pool of differently sized machines, where the question is how much headroom each one has rather than how loaded it is.

The controller keeps the last weight it computed successfully for that member. A member that is briefly unreachable therefore holds its share rather than dropping to the floor and taking a thundering share back when it returns.

The loop is restarted on every configuration reload.

healthcheck:
path: /health
method: GET
status: 200
interval: 10s
timeout: 3s
scheme: https
port: "7102"
hostname: api.internal
headers: { X-Probe: gateway }

A health check can be set on the default profile, so it applies everywhere, or per profile. Failover depends on one; the other algorithms use it to take a bad server out of rotation.

The whole reload is rejected, and the previous configuration stays live, when:

  • a profile names a backend that is not declared
  • profiles and backends form a cycle
  • weighted or adaptive declares no members
  • mirror has no primary, or failover has no primary and fallback
  • a weight is negative, or a mirror percent is out of range
  • adaptive has no reconcile block, an unknown formula, min_weight below 1, or a max_weight below min_weight
  • reconcile.interval is under 5s

The error names the profile and the offending field.