Rate Limiting Incoming Requests in ASGI Apps¶
Most rate-limiting advice for asyncio is about being a polite client. A service also needs to protect itself from clients: a misconfigured integration that retries in a tight loop, a script hammering one endpoint, a tenant whose traffic spikes tenfold. The usual tool is a token bucket per client key, enforced in ASGI middleware before any application code runs. Measured with a pure ASGI middleware in front of a Starlette app, a burst of 200 requests from one API key within 18 ms let 20 through (the bucket's burst size) and answered 180 with 429 and Retry-After: 1; a different key making a request in the middle of that burst got 200; and a client pacing itself at 40 requests per second against a 50 per second limit had all 80 requests accepted. This guide builds that middleware and covers the deployment details that decide whether the limit you configured is the limit you get.
Prerequisites¶
- Python 3.11+, any ASGI framework (Starlette and FastAPI shown);
pip install httpxfor the tests. - Token bucket mechanics, from token bucket rate limiter for asyncio clients.
- Pure ASGI middleware, from writing pure ASGI middleware.
1. Write a token bucket that never sleeps¶
Server-side limiting must answer immediately — allow or reject — never wait. The bucket refills lazily from elapsed time on each check:
import time
class TokenBucket:
__slots__ = ("rate", "burst", "tokens", "updated")
def __init__(self, rate: float, burst: int) -> None:
self.rate, self.burst = rate, burst
self.tokens = float(burst)
self.updated = time.monotonic()
def allow(self, cost: float = 1.0) -> tuple[bool, float]:
now = time.monotonic()
self.tokens = min(self.burst, self.tokens + (now - self.updated) * self.rate)
self.updated = now
if self.tokens >= cost:
self.tokens -= cost
return True, 0.0
return False, (cost - self.tokens) / self.rate # seconds until it would pass
allow() is synchronous and does no I/O, so on a single event loop it is atomic: no other request can interleave between reading and updating tokens. rate is the sustained requests per second and burst the size of a spike tolerated without delay. Returning the wait time lets the response carry an accurate Retry-After.
Verify: call allow() burst + 5 times in a tight loop; exactly burst succeed.
2. Enforce it in pure ASGI middleware¶
Middleware written against the raw ASGI interface runs for every request with almost no overhead and works with any framework:
from starlette.responses import PlainTextResponse
class RateLimitMiddleware:
def __init__(self, app, rate: float = 50, burst: int = 20) -> None:
self.app, self.rate, self.burst = app, rate, burst
self.buckets: dict[bytes, TokenBucket] = {}
async def __call__(self, scope, receive, send):
if scope["type"] != "http":
return await self.app(scope, receive, send)
key = self.client_key(scope)
bucket = self.buckets.get(key)
if bucket is None:
bucket = self.buckets[key] = TokenBucket(self.rate, self.burst)
ok, wait = bucket.allow()
if not ok:
response = PlainTextResponse(
"rate limited", status_code=429,
headers={"Retry-After": str(max(1, round(wait)))},
)
return await response(scope, receive, send)
await self.app(scope, receive, send)
@staticmethod
def client_key(scope) -> bytes:
headers = dict(scope["headers"])
return headers.get(b"x-api-key") or (scope.get("client") or ("anon",))[0].encode()
Measured: 200 concurrent requests from key A — 20 accepted, 180 rejected with Retry-After: 1; key B, during that burst, accepted. Rejected requests never reach the application, so they cost a dictionary lookup and a tiny response rather than a database query. Keying on the API key when there is one and the client address otherwise is common; behind a proxy, the address in scope["client"] is the proxy's, so take the forwarded address only from a trusted proxy header.
Verify: an integration test that bursts beyond burst gets 429s with Retry-After, and a well-behaved client below rate never does.
3. Bound the bucket table¶
One bucket per key, kept forever, is a memory leak keyed by every client address that ever connected — trivially exploitable with spoofed or rotating addresses. Evict buckets that are full and idle, since a full bucket behaves exactly like a missing one:
async def sweep(self, every: float = 30.0, idle: float = 60.0) -> None:
while True:
await asyncio.sleep(every)
now = time.monotonic()
stale = [k for k, b in self.buckets.items()
if now - b.updated > idle and b.tokens + (now - b.updated) * b.rate >= b.burst]
for k in stale:
del self.buckets[k]
Start sweep() from the application's lifespan, as described in managing startup and shutdown with ASGI lifespan. Also cap the table size outright: if it exceeds a ceiling, reject new keys or fall back to a shared bucket, so a flood of unique keys cannot exhaust memory between sweeps.
Verify: after traffic from 100,000 distinct addresses stops, the table returns to near zero within two sweep intervals.
4. Account for multiple workers and replicas¶
Each worker process has its own middleware instance and its own buckets. With rate=50 and four uvicorn workers, a client spread across connections gets up to 200 per second; across three replicas, 600. Two ways to handle it:
- Divide the limit by the number of processes when an approximate limit is acceptable. Load balancers rarely spread one client evenly, so the effective limit varies.
- Keep the counter in Redis when the limit is a contract — a paid tier's quota, an abuse threshold. An atomic Lua script per request costs a round trip, typically well under a millisecond on a local network:
async def allow_global(redis, key: str, rate: int, window_s: int = 1) -> bool:
bucket = f"rl:{key}:{int(time.time() // window_s)}"
count = await redis.incr(bucket)
if count == 1:
await redis.expire(bucket, window_s * 2)
return count <= rate
The fixed-window version above is the simplest; its edge-of-window burst and the sliding-window fix are covered in sliding window rate limiting with Redis and asyncio.
Verify: with all workers running, a load test from one key never exceeds the configured global rate.
5. Exempt and prioritise deliberately¶
Not every request should share the limit. Health checks from the orchestrator, internal service-to-service calls and webhooks you depend on should bypass client limits, and expensive endpoints may need a lower limit than cheap ones:
ROUTE_COSTS = {"/search": 5.0, "/export": 20.0}
EXEMPT = {"/healthz", "/readyz"}
async def __call__(self, scope, receive, send):
if scope["type"] != "http" or scope["path"] in EXEMPT:
return await self.app(scope, receive, send)
cost = ROUTE_COSTS.get(scope["path"], 1.0)
ok, wait = self.bucket_for(scope).allow(cost)
...
Weighting by cost lets one bucket protect both cheap and expensive endpoints. Rejecting health checks is a classic self-inflicted outage: the orchestrator sees failures, restarts healthy pods, and the remaining pods receive even more load. For load that threatens the whole service rather than one client, a global shedder in front of everything is the right tool, as in load shedding when the event loop is overloaded.
Verify: health checks succeed while a client is being rate limited, and one /export call consumes twenty tokens.
Verification¶
Request limiting works when:
- Bursts beyond
burstget 429 with an accurateRetry-After, and paced clients belowrateare never limited. - One client's burst does not affect other keys.
- The bucket table is bounded by sweeping and a hard cap.
- The effective limit is known across workers and replicas, or enforced in a shared store.
Diagnostic Hook: export 429s per key (top N only, to bound cardinality) and per route, plus the bucket table size. A single key dominating the 429s is a misbehaving client to contact; 429s spread across many keys on one route usually mean that route's cost is set too high or its limit too low for legitimate traffic.
Pitfalls & edge cases¶
- Keying on the proxy address. Every client behind the load balancer shares one bucket.
- Unbounded bucket tables. A memory leak keyed by attacker-chosen addresses.
- Per-process limits treated as global. The real limit is the setting times the number of processes.
- Rate limiting health checks. The orchestrator restarts healthy instances.
Frequently Asked Questions¶
How do I rate limit requests in FastAPI or Starlette?
Add pure ASGI middleware that keeps a token bucket per client key, checks it before calling the app, and returns 429 with a Retry-After header when the bucket is empty. In testing, a 200-request burst against burst 20 let 20 through.
Should a server-side rate limiter wait or reject?
Reject immediately with 429 and Retry-After. Waiting holds connections and memory for requests the client may abandon, and lets a misbehaving client consume server capacity.
Why is my rate limit higher than configured?
Each worker process keeps its own buckets, so the effective limit is the setting multiplied by workers and replicas. Divide the setting, or store counters in Redis for an exact global limit.
How do I stop the rate limiter's memory growing?
Periodically remove buckets that are idle and full, which behave like missing ones, and cap the total number of buckets so a flood of unique keys cannot exhaust memory.
Related¶
- Rate Limiting & Throttling — up to the topic overview.
- Per-tenant rate limits in async services — limits that follow a customer, not a connection.
- Concurrent Execution & Worker Patterns — the section overview.