Per-Tenant Rate Limits in Async Services¶
In a multi-tenant service, global limits protect the service and do nothing for its customers. A single tenant running a bulk import can fill every worker slot, every connection and every queue position, and the other tenants — who did nothing unusual — see their latency explode. Measured in a simulation with a shared pool of 10 concurrent slots: when a noisy tenant submitted 500 requests of 20 ms and a quiet tenant then submitted 10, the quiet tenant's median latency was 982 ms, because it queued behind the entire flood. Capping each tenant at 6 of the 10 slots brought the quiet tenant to 30 ms median and 51 ms worst case — at the price of the noisy tenant's batch taking 1.70 s instead of 1.03 s. This guide combines per-tenant rate limits and concurrency shares, and shows how to give back the capacity the cap leaves idle.
Prerequisites¶
- Python 3.11+, stdlib only; the shared-counter variant uses Redis.
- Request limiting middleware, from rate limiting incoming requests in ASGI apps.
- Fair scheduling, from fair scheduling across tenants in a worker pool.
1. Identify the tenant early, and key everything on it¶
Every limit in this guide is keyed by tenant, so resolve the tenant once, as early as possible, and carry it in a context variable for the rest of the request:
import contextvars
tenant_var: contextvars.ContextVar[str] = contextvars.ContextVar("tenant", default="anonymous")
class TenantMiddleware:
def __init__(self, app, resolve) -> None:
self.app, self.resolve = app, resolve
async def __call__(self, scope, receive, send):
if scope["type"] != "http":
return await self.app(scope, receive, send)
token = tenant_var.set(await self.resolve(scope)) # from API key, JWT claim, subdomain
try:
await self.app(scope, receive, send)
finally:
tenant_var.reset(token)
Resolve from something the tenant cannot forge — a validated API key or token claim — rather than a header the client sets. Downstream code, including outbound clients and background jobs, can then read tenant_var to apply the tenant's limits without passing the tenant around, as described in propagating request IDs with contextvars.
Verify: every log line and metric for a request carries the tenant id.
2. Give each tenant a rate according to its tier¶
Rates come from the tenant's plan, not from a global constant. A token bucket per tenant, parameterised by tier:
import time
from dataclasses import dataclass
@dataclass(frozen=True)
class Tier:
rate: float # sustained requests per second
burst: int # spike tolerated without throttling
share: int # concurrent request slots (step 3)
TIERS = {"free": Tier(2, 10, 2), "pro": Tier(20, 60, 6), "enterprise": Tier(100, 300, 20)}
class TenantBuckets:
def __init__(self, tier_of) -> None:
self.tier_of = tier_of
self.buckets: dict[str, TokenBucket] = {}
def allow(self, tenant: str) -> tuple[bool, float]:
b = self.buckets.get(tenant)
if b is None:
t = TIERS[self.tier_of(tenant)]
b = self.buckets[tenant] = TokenBucket(t.rate, t.burst)
return b.allow()
TokenBucket is the non-blocking bucket from the ASGI rate-limiting guide. Return 429 with Retry-After when allow() fails, and include the tier in the response so the client knows which limit it hit. When a tenant upgrades, drop its bucket so the next request builds one with the new parameters. Rates protect against sustained floods; they do not protect against a tenant whose requests are each slow, which is what step 3 is for.
Verify: a free-tier tenant bursting beyond 10 gets 429s while an enterprise tenant at the same moment does not.
3. Cap each tenant's share of concurrency¶
Rate limits count arrivals; the resource that actually runs out is concurrent work in flight. Give each tenant a semaphore sized by tier, acquired before the global pool:
import asyncio
from collections import defaultdict
global_slots = asyncio.Semaphore(10)
tenant_slots: dict[str, asyncio.Semaphore] = {}
def slots_for(tenant: str) -> asyncio.Semaphore:
s = tenant_slots.get(tenant)
if s is None:
s = tenant_slots[tenant] = asyncio.Semaphore(TIERS[tier_of(tenant)].share)
return s
async def handle(request):
tenant = tenant_var.get()
async with slots_for(tenant): # the tenant's share first
async with global_slots: # then the service's capacity
return await do_work(request)
Order matters: acquiring the tenant's semaphore first means a flood queues on its own semaphore, leaving global slots for others. Measured: with a share of 6 out of 10, the quiet tenant's median latency fell from 982 ms to 30 ms. Keep the sum of shares for typical concurrent tenants comfortably above the global capacity, or the caps will leave the service idle while tenants wait. This is the bulkhead idea from isolating tenants with bulkheads applied to request handling.
Verify: under a flood from one tenant, another tenant's p99 latency stays near its unloaded value.
4. Give idle capacity back to busy tenants¶
A hard cap wastes capacity when only one tenant is active: the noisy tenant's batch took 1.70 s instead of 1.03 s because 4 of the 10 slots sat idle. A work-conserving variant lets a tenant exceed its share only when nobody else is waiting:
class FairSlots:
def __init__(self, capacity: int, share: int) -> None:
self.capacity, self.share = capacity, share
self.in_use: dict[str, int] = defaultdict(int)
self.waiting: dict[str, int] = defaultdict(int) # waiters per tenant, not a set
self.cond = asyncio.Condition()
def _can_run(self, tenant: str) -> bool:
if sum(self.in_use.values()) >= self.capacity:
return False
if self.in_use[tenant] < self.share:
return True # within its share: always
others = any(n for t, n in self.waiting.items() if t != tenant)
return not others # above it: only if nobody else waits
async def acquire(self, tenant: str) -> None:
async with self.cond:
self.waiting[tenant] += 1
try:
await self.cond.wait_for(lambda: self._can_run(tenant))
finally:
self.waiting[tenant] -= 1
self.in_use[tenant] += 1
async def release(self, tenant: str) -> None:
async with self.cond:
self.in_use[tenant] -= 1
self.cond.notify_all()
A tenant within its share always proceeds when there is capacity; above its share, it proceeds only if no other tenant is waiting. Measured on the same flood: the quiet tenant's median latency was 57 ms (79 ms worst case) and the noisy batch finished in 1.08 s — nearly the uncapped 1.03 s, against 1.70 s with the hard cap. Count waiters per tenant, not in a set: a first draft that tracked waiting tenants in a set removed a tenant as soon as one of its tasks was admitted, so the noisy tenant looked unopposed and the quiet tenant still waited about a second. notify_all() wakes every waiter on each release, which is fine for hundreds of waiters; for thousands, use per-tenant queues as in the fair-scheduling guide linked above.
Verify: a lone tenant uses the whole capacity; when a second arrives, the first's in-flight count falls to its share within one request duration.
5. Share tenant limits across replicas¶
Rates and shares enforced per process multiply by the number of processes, and tenants are rarely spread evenly across them. For plan-level quotas, keep counters in Redis keyed by tenant — the multi-window Lua script from enforcing multiple rate limits at once works unchanged with tenant-scoped keys:
async def allow_tenant(r, tenant: str) -> bool:
t = TIERS[tier_of(tenant)]
keys = [f"rl:{tenant}:sec", f"rl:{tenant}:day"]
... # same check-all-then-record script, limits from the tier
Keep the concurrency share local — it protects each process's event loop and pools, which are local resources — and move only the contractual rate quotas to the shared store. That costs one Redis round trip per request for the quota check and none for the share.
Verify: a tenant's daily usage reported by Redis matches the sum across all replicas' access logs.
Verification¶
Tenant isolation works when:
- Every limit is keyed by a tenant id resolved from authenticated data.
- Rates follow the tenant's tier, and 429 responses name the limit that was hit.
- A flood from one tenant leaves others' latency near normal.
- Idle capacity is not wasted when only one tenant is busy, if throughput matters.
Diagnostic Hook: export per-tenant in-flight count, wait time for a tenant slot, and 429s — bounded to the top tenants by traffic to control cardinality. A tenant whose slot wait is high while global capacity sits idle means its share is too small for its tier; a tenant whose 429 rate is high while others are unaffected is the isolation working as designed.
Pitfalls & edge cases¶
- Tenant from an unvalidated header. A client can impersonate another tenant's quota.
- Global slot acquired before the tenant slot. Floods occupy shared capacity while waiting on their own share.
- Shares summing below capacity. The service idles while tenants queue.
- Unbounded per-tenant tables. Evict buckets and semaphores of inactive tenants.
Frequently Asked Questions¶
How do I stop one tenant slowing down every other tenant?
Cap each tenant's concurrent requests with a per-tenant semaphore acquired before the shared pool. In a simulation, a cap of 6 of 10 slots took a quiet tenant's median latency during another tenant's flood from 982 ms to 30 ms.
Should per-tenant limits be rates or concurrency caps?
Both. Rate limits enforce plan quotas on arrivals; concurrency caps protect latency when one tenant's requests are slow or numerous. They fail in different ways and complement each other.
Don't per-tenant caps waste capacity?
Hard caps do when only one tenant is busy. A work-conserving rule lets a tenant exceed its share only when no other tenant is waiting, keeping isolation without idle slots.
How do I enforce tenant quotas across several replicas?
Keep the quota counters in Redis keyed by tenant and check them atomically with a Lua script. Keep concurrency shares per process, since they protect local resources.
Related¶
- Rate Limiting & Throttling — up to the topic overview.
- Limiting concurrency per job type — the same isolation for background work.
- Concurrent Execution & Worker Patterns — the section overview.