Skip to content

Rate Limiting Incoming Requests in ASGI Apps

Most rate-limiting advice for asyncio is about being a polite client. A service also needs to protect itself from clients: a misconfigured integration that retries in a tight loop, a script hammering one endpoint, a tenant whose traffic spikes tenfold. The usual tool is a token bucket per client key, enforced in ASGI middleware before any application code runs. Measured with a pure ASGI middleware in front of a Starlette app, a burst of 200 requests from one API key within 18 ms let 20 through (the bucket's burst size) and answered 180 with 429 and Retry-After: 1; a different key making a request in the middle of that burst got 200; and a client pacing itself at 40 requests per second against a 50 per second limit had all 80 requests accepted. This guide builds that middleware and covers the deployment details that decide whether the limit you configured is the limit you get.

Prerequisites

1. Write a token bucket that never sleeps

Server-side limiting must answer immediately — allow or reject — never wait. The bucket refills lazily from elapsed time on each check:

import time


class TokenBucket:
    __slots__ = ("rate", "burst", "tokens", "updated")

    def __init__(self, rate: float, burst: int) -> None:
        self.rate, self.burst = rate, burst
        self.tokens = float(burst)
        self.updated = time.monotonic()

    def allow(self, cost: float = 1.0) -> tuple[bool, float]:
        now = time.monotonic()
        self.tokens = min(self.burst, self.tokens + (now - self.updated) * self.rate)
        self.updated = now
        if self.tokens >= cost:
            self.tokens -= cost
            return True, 0.0
        return False, (cost - self.tokens) / self.rate        # seconds until it would pass

allow() is synchronous and does no I/O, so on a single event loop it is atomic: no other request can interleave between reading and updating tokens. rate is the sustained requests per second and burst the size of a spike tolerated without delay. Returning the wait time lets the response carry an accurate Retry-After.

Verify: call allow() burst + 5 times in a tight loop; exactly burst succeed.

2. Enforce it in pure ASGI middleware

Middleware written against the raw ASGI interface runs for every request with almost no overhead and works with any framework:

from starlette.responses import PlainTextResponse


class RateLimitMiddleware:
    def __init__(self, app, rate: float = 50, burst: int = 20) -> None:
        self.app, self.rate, self.burst = app, rate, burst
        self.buckets: dict[bytes, TokenBucket] = {}

    async def __call__(self, scope, receive, send):
        if scope["type"] != "http":
            return await self.app(scope, receive, send)
        key = self.client_key(scope)
        bucket = self.buckets.get(key)
        if bucket is None:
            bucket = self.buckets[key] = TokenBucket(self.rate, self.burst)
        ok, wait = bucket.allow()
        if not ok:
            response = PlainTextResponse(
                "rate limited", status_code=429,
                headers={"Retry-After": str(max(1, round(wait)))},
            )
            return await response(scope, receive, send)
        await self.app(scope, receive, send)

    @staticmethod
    def client_key(scope) -> bytes:
        headers = dict(scope["headers"])
        return headers.get(b"x-api-key") or (scope.get("client") or ("anon",))[0].encode()

Measured: 200 concurrent requests from key A — 20 accepted, 180 rejected with Retry-After: 1; key B, during that burst, accepted. Rejected requests never reach the application, so they cost a dictionary lookup and a tiny response rather than a database query. Keying on the API key when there is one and the client address otherwise is common; behind a proxy, the address in scope["client"] is the proxy's, so take the forwarded address only from a trusted proxy header.

Verify: an integration test that bursts beyond burst gets 429s with Retry-After, and a well-behaved client below rate never does.

200 requests from one key in 18 ms 2 horizontal bars comparing rejected with 429 with the others. 200 requests from one key in 18 ms rejected with 429 180, Retry-After: 1 allowed 20, the burst size Pure ASGI middleware in front of Starlette; a second key during the burst was allowed. The burst size, not the rate, decides how much of a sudden spike gets through.

3. Bound the bucket table

One bucket per key, kept forever, is a memory leak keyed by every client address that ever connected — trivially exploitable with spoofed or rotating addresses. Evict buckets that are full and idle, since a full bucket behaves exactly like a missing one:

    async def sweep(self, every: float = 30.0, idle: float = 60.0) -> None:
        while True:
            await asyncio.sleep(every)
            now = time.monotonic()
            stale = [k for k, b in self.buckets.items()
                     if now - b.updated > idle and b.tokens + (now - b.updated) * b.rate >= b.burst]
            for k in stale:
                del self.buckets[k]

Start sweep() from the application's lifespan, as described in managing startup and shutdown with ASGI lifespan. Also cap the table size outright: if it exceeds a ceiling, reject new keys or fall back to a shared bucket, so a flood of unique keys cannot exhaust memory between sweeps.

Verify: after traffic from 100,000 distinct addresses stops, the table returns to near zero within two sweep intervals.

4. Account for multiple workers and replicas

Each worker process has its own middleware instance and its own buckets. With rate=50 and four uvicorn workers, a client spread across connections gets up to 200 per second; across three replicas, 600. Two ways to handle it:

  • Divide the limit by the number of processes when an approximate limit is acceptable. Load balancers rarely spread one client evenly, so the effective limit varies.
  • Keep the counter in Redis when the limit is a contract — a paid tier's quota, an abuse threshold. An atomic Lua script per request costs a round trip, typically well under a millisecond on a local network:
async def allow_global(redis, key: str, rate: int, window_s: int = 1) -> bool:
    bucket = f"rl:{key}:{int(time.time() // window_s)}"
    count = await redis.incr(bucket)
    if count == 1:
        await redis.expire(bucket, window_s * 2)
    return count <= rate

The fixed-window version above is the simplest; its edge-of-window burst and the sliding-window fix are covered in sliding window rate limiting with Redis and asyncio.

Verify: with all workers running, a load test from one key never exceeds the configured global rate.

Where should the counter live? A grid of 4 rows by 3 columns. Where should the counter live? property in-process bucket Redis counter limit across 4 workers up to 4x the setting exact cost per request microseconds one Redis round trip if the store fails n/a fail open or closed: choose use for abuse and burst protection quotas and paid tiers In-process limits are cheap guards; contractual limits need a shared counter.

5. Exempt and prioritise deliberately

Not every request should share the limit. Health checks from the orchestrator, internal service-to-service calls and webhooks you depend on should bypass client limits, and expensive endpoints may need a lower limit than cheap ones:

ROUTE_COSTS = {"/search": 5.0, "/export": 20.0}
EXEMPT = {"/healthz", "/readyz"}


async def __call__(self, scope, receive, send):
    if scope["type"] != "http" or scope["path"] in EXEMPT:
        return await self.app(scope, receive, send)
    cost = ROUTE_COSTS.get(scope["path"], 1.0)
    ok, wait = self.bucket_for(scope).allow(cost)
    ...

Weighting by cost lets one bucket protect both cheap and expensive endpoints. Rejecting health checks is a classic self-inflicted outage: the orchestrator sees failures, restarts healthy pods, and the remaining pods receive even more load. For load that threatens the whole service rather than one client, a global shedder in front of everything is the right tool, as in load shedding when the event loop is overloaded.

Verify: health checks succeed while a client is being rate limited, and one /export call consumes twenty tokens.

What the middleware does with each request A flow of 4 stages. What the middleware does with each request exempt path? health checks pass bucket for key API key or client IP allow(route cost) expensive routes cost more app or 429 Retry-After from wait Every decision happens before the application sees the request.

Verification

Request limiting works when:

  • Bursts beyond burst get 429 with an accurate Retry-After, and paced clients below rate are never limited.
  • One client's burst does not affect other keys.
  • The bucket table is bounded by sweeping and a hard cap.
  • The effective limit is known across workers and replicas, or enforced in a shared store.

Diagnostic Hook: export 429s per key (top N only, to bound cardinality) and per route, plus the bucket table size. A single key dominating the 429s is a misbehaving client to contact; 429s spread across many keys on one route usually mean that route's cost is set too high or its limit too low for legitimate traffic.

Pitfalls & edge cases

  • Keying on the proxy address. Every client behind the load balancer shares one bucket.
  • Unbounded bucket tables. A memory leak keyed by attacker-chosen addresses.
  • Per-process limits treated as global. The real limit is the setting times the number of processes.
  • Rate limiting health checks. The orchestrator restarts healthy instances.

Frequently Asked Questions

How do I rate limit requests in FastAPI or Starlette?

Add pure ASGI middleware that keeps a token bucket per client key, checks it before calling the app, and returns 429 with a Retry-After header when the bucket is empty. In testing, a 200-request burst against burst 20 let 20 through.

Should a server-side rate limiter wait or reject?

Reject immediately with 429 and Retry-After. Waiting holds connections and memory for requests the client may abandon, and lets a misbehaving client consume server capacity.

Why is my rate limit higher than configured?

Each worker process keeps its own buckets, so the effective limit is the setting multiplied by workers and replicas. Divide the setting, or store counters in Redis for an exact global limit.

How do I stop the rate limiter's memory growing?

Periodically remove buckets that are idle and full, which behave like missing ones, and cap the total number of buckets so a flood of unique keys cannot exhaust memory.