Skip to content

Enforcing Multiple Rate Limits at Once

Real APIs rarely have one limit. A typical contract says "10 requests per second, 500 per minute, 10,000 per day", sometimes plus a concurrency cap. Stacking one limiter per rule looks straightforward — acquire the per-second limiter, then the per-minute one — and is subtly wrong. Each limiter records the call when it lets it through, but the call is only sent after the last limiter allows it, so earlier limiters record times that are too early and let later calls through too soon. In a test with limits of 5 per second and 12 per 3 seconds, sequential acquisition produced 2 violations of the 5-per-second rule over 30 calls; a combined limiter that asked every rule for its wait time, slept for the longest, and then recorded the call in all rules at once produced 0, in the same 7.0 s. This guide builds the combined limiter and extends it to concurrency caps and distributed counters.

Prerequisites

1. Make each limit answer "how long must I wait?"

Instead of each limiter blocking on its own, give every limit the same two operations: report the wait required right now, and record a call. A sliding log does both exactly:

import collections
import time


class SlidingLog:
    def __init__(self, limit: int, window: float) -> None:
        self.limit, self.window = limit, window
        self.log: collections.deque[float] = collections.deque()

    def _trim(self, now: float) -> None:
        while self.log and now - self.log[0] >= self.window:
            self.log.popleft()

    def wait_time(self, now: float) -> float:
        self._trim(now)
        if len(self.log) < self.limit:
            return 0.0
        return self.log[0] + self.window - now        # until the oldest call leaves the window

    def record(self, now: float) -> None:
        self.log.append(now)

A sliding log uses memory proportional to the limit — 500 timestamps for a per-minute limit of 500 — which is fine for client-side limits. For limits in the tens of thousands, a sliding-window counter approximation is cheaper; the interface stays the same.

Verify: with limit calls recorded inside the window, wait_time() returns exactly the time until the oldest one expires.

2. Check every limit, then record once

The combined limiter asks all limits for their wait, sleeps for the maximum, re-checks (other callers may have moved first), and only when every limit reports zero records the call in all of them at the same instant:

import asyncio


class MultiLimit:
    def __init__(self, *limits: SlidingLog) -> None:
        self.limits = limits
        self._lock = asyncio.Lock()                   # one caller evaluates at a time

    async def acquire(self) -> None:
        async with self._lock:
            while True:
                now = time.monotonic()
                wait = max(l.wait_time(now) for l in self.limits)
                if wait <= 0:
                    for l in self.limits:
                        l.record(now)                 # same timestamp everywhere
                    return
                await asyncio.sleep(wait)


api_limits = MultiLimit(SlidingLog(10, 1.0), SlidingLog(500, 60.0), SlidingLog(10_000, 86_400.0))


async def call_api(session, url):
    await api_limits.acquire()
    async with session.get(url) as r:
        return await r.json()

Measured with 5 per second and 12 per 3 seconds over 30 calls: zero violations of either rule. The lock serialises admission, not the calls themselves — once acquire() returns, callers run concurrently. Holding the lock while sleeping is deliberate: it keeps callers in FIFO order and stops a stampede of re-checks when a window opens.

Verify: record send timestamps under load and check every window against every limit; there should be no violations.

How the combined limiter admits one call A flow of 4 stages. How the combined limiter admits one call ask every limit wait_time(now) sleep the maximum longest wait wins re-check all others may have moved record in all at once same timestamp Recording only after every rule agrees is what keeps the timestamps honest.

3. See why sequential acquisition violates limits

The tempting composition acquires limiters one after another:

async def sequential(limits):
    for limit in limits:
        while (w := limit.wait_time(time.monotonic())) > 0:
            await asyncio.sleep(w)
        limit.record(time.monotonic())                 # recorded now, sent later

When the per-second limiter admits the call and records it at time t, but the per-3-second limiter then makes it wait until t + 0.8 s, the per-second log believes a call happened at t that actually happened at t + 0.8. A second later, that phantom entry expires early and lets an extra call through — which is how sequential acquisition produced 2 violations of the 5-per-second rule in 30 calls. The reverse failure also exists: a call admitted by the first limiter and then cancelled while waiting on the second has consumed capacity it never used.

Verify: run the violation check against sequential acquisition with your real limits; any non-zero result is this bug.

Violations over 30 calls with two limits 2 horizontal bars comparing sequential acquire with the others. Violations over 30 calls with two limits sequential acquire 2 violations of 5 per 1 s check all, record once 0 violations Both runs took 7.0 s; the difference is entirely in when calls are recorded. Recording a call before it is actually allowed by every rule corrupts the earlier rules.

4. Add a concurrency cap alongside the rates

Many APIs also limit simultaneous requests. A concurrency cap is not a rate — it is released when the call finishes — so combine it as a separate stage, acquired after the rate limits and held for the call's duration:

class ApiGate:
    def __init__(self, rates: MultiLimit, max_in_flight: int) -> None:
        self.rates = rates
        self.slots = asyncio.Semaphore(max_in_flight)

    async def __aenter__(self):
        await self.slots.acquire()                     # hold a slot for the whole call
        try:
            await self.rates.acquire()                 # then wait for every rate rule
        except BaseException:
            self.slots.release()
            raise
        return self

    async def __aexit__(self, *exc):
        self.slots.release()


gate = ApiGate(api_limits, max_in_flight=4)


async def call_api(session, url):
    async with gate:
        async with session.get(url) as r:
            return await r.json()

Acquiring the slot first means a caller only consumes rate capacity when it can actually send, avoiding the "admitted by the rate, then stuck waiting for a slot" skew from step 3. The semaphore pattern itself is covered in limiting concurrent requests with asyncio.Semaphore.

Verify: under load, in-flight calls never exceed the cap and send timestamps satisfy every rate rule.

5. Coordinate limits across processes

When several workers share one API key, each process's limiter only sees its own calls. Move the check-and-record into Redis as a single atomic script over all windows:

MULTI = """
local now = tonumber(ARGV[1])
for i = 1, #KEYS do
  local limit, window = tonumber(ARGV[i*2]), tonumber(ARGV[i*2+1])
  redis.call('ZREMRANGEBYSCORE', KEYS[i], '-inf', now - window)
  if redis.call('ZCARD', KEYS[i]) >= limit then
    local oldest = redis.call('ZRANGE', KEYS[i], 0, 0, 'WITHSCORES')[2]
    return tostring(tonumber(oldest) + window - now)          -- wait required
  end
end
for i = 1, #KEYS do
  redis.call('ZADD', KEYS[i], now, ARGV[#ARGV])
  redis.call('EXPIRE', KEYS[i], math.ceil(tonumber(ARGV[i*2+1])))
end
return '0'
"""


async def acquire_global(r, member: str) -> None:
    rules = [("rl:sec", 10, 1), ("rl:min", 500, 60)]
    keys = [k for k, _, _ in rules]
    while True:
        args = [time.time()] + [x for _, lim, win in rules for x in (lim, win)] + [member]
        wait = float(await r.eval(MULTI, len(keys), *keys, *args))
        if wait <= 0:
            return
        await asyncio.sleep(wait)

Lua scripts run atomically in Redis, so "check every rule, then record in every rule" holds across the whole fleet. Use a unique member per call (a UUID) so two calls at the same timestamp both count. This is the multi-window version of the single-limit script in the sliding-window guide linked above.

Verify: with four workers sharing the key, the API's own rate-limit headers never report an exceeded limit.

How many limiters, and where? A decision on Who shares the API's limits with 3 outcomes. How many limiters, and where? Who shares the API's limits? one process MultiLimit in memory check all, record once many processes, one key Redis Lua script atomic across the fleet plus a concurrency cap add a Semaphore stage slot first, then rates Whatever the scope, never record a call in one rule before every rule has admitted it.

Verification

Multi-limit enforcement is correct when:

  • Every rule is checked before any rule records, and all record the same timestamp.
  • Measured send times satisfy every window under load, with zero violations.
  • Concurrency caps are a separate stage acquired before the rate rules.
  • Shared keys use an atomic shared check across processes.

Diagnostic Hook: export, per rule, the current usage as a fraction of its limit and the time callers spend waiting on it. The rule with the highest usage is the binding constraint; if the per-day rule binds by mid-afternoon, no amount of per-second tuning will help, and the conversation with the provider should be about the daily quota.

Pitfalls & edge cases

  • Acquiring limiters one after another. Earlier rules record too early; measured violations follow.
  • Cancellation between stages. Capacity consumed by a cancelled caller is not reclaimed in sequential designs.
  • Per-process limiters for a shared key. The provider sees the sum.
  • Clock differences in distributed limits. Use Redis server time, or keep hosts synchronised.

Frequently Asked Questions

How do I enforce a per-second and a per-minute rate limit together?

Ask every limit how long the caller must wait, sleep for the longest, re-check, and when every limit allows the call, record it in all of them with the same timestamp. In testing this produced no violations, while acquiring limiters one after another produced two.

Why does chaining rate limiters cause violations?

The first limiter records the call when it admits it, but the call is only sent after the later limiters allow it. The early timestamp expires too soon and lets an extra call through later.

How do I combine a rate limit with a concurrency limit?

Acquire the concurrency slot first and hold it for the call, then wait for the rate rules, so rate capacity is only consumed by callers that can actually send.

How do I share multiple rate limits across processes?

Run the check-all-then-record-all logic as a single Lua script in Redis over sorted sets, one per rule. Scripts execute atomically, so the rules hold across every process.