Enforcing Multiple Rate Limits at Once¶
Real APIs rarely have one limit. A typical contract says "10 requests per second, 500 per minute, 10,000 per day", sometimes plus a concurrency cap. Stacking one limiter per rule looks straightforward — acquire the per-second limiter, then the per-minute one — and is subtly wrong. Each limiter records the call when it lets it through, but the call is only sent after the last limiter allows it, so earlier limiters record times that are too early and let later calls through too soon. In a test with limits of 5 per second and 12 per 3 seconds, sequential acquisition produced 2 violations of the 5-per-second rule over 30 calls; a combined limiter that asked every rule for its wait time, slept for the longest, and then recorded the call in all rules at once produced 0, in the same 7.0 s. This guide builds the combined limiter and extends it to concurrency caps and distributed counters.
Prerequisites¶
- Python 3.11+, stdlib only; the distributed variant uses
redis.asyncio. - Sliding windows, from sliding window rate limiting with Redis and asyncio.
- Token and leaky buckets, from smoothing bursts with a leaky bucket.
1. Make each limit answer "how long must I wait?"¶
Instead of each limiter blocking on its own, give every limit the same two operations: report the wait required right now, and record a call. A sliding log does both exactly:
import collections
import time
class SlidingLog:
def __init__(self, limit: int, window: float) -> None:
self.limit, self.window = limit, window
self.log: collections.deque[float] = collections.deque()
def _trim(self, now: float) -> None:
while self.log and now - self.log[0] >= self.window:
self.log.popleft()
def wait_time(self, now: float) -> float:
self._trim(now)
if len(self.log) < self.limit:
return 0.0
return self.log[0] + self.window - now # until the oldest call leaves the window
def record(self, now: float) -> None:
self.log.append(now)
A sliding log uses memory proportional to the limit — 500 timestamps for a per-minute limit of 500 — which is fine for client-side limits. For limits in the tens of thousands, a sliding-window counter approximation is cheaper; the interface stays the same.
Verify: with limit calls recorded inside the window, wait_time() returns exactly the time until the oldest one expires.
2. Check every limit, then record once¶
The combined limiter asks all limits for their wait, sleeps for the maximum, re-checks (other callers may have moved first), and only when every limit reports zero records the call in all of them at the same instant:
import asyncio
class MultiLimit:
def __init__(self, *limits: SlidingLog) -> None:
self.limits = limits
self._lock = asyncio.Lock() # one caller evaluates at a time
async def acquire(self) -> None:
async with self._lock:
while True:
now = time.monotonic()
wait = max(l.wait_time(now) for l in self.limits)
if wait <= 0:
for l in self.limits:
l.record(now) # same timestamp everywhere
return
await asyncio.sleep(wait)
api_limits = MultiLimit(SlidingLog(10, 1.0), SlidingLog(500, 60.0), SlidingLog(10_000, 86_400.0))
async def call_api(session, url):
await api_limits.acquire()
async with session.get(url) as r:
return await r.json()
Measured with 5 per second and 12 per 3 seconds over 30 calls: zero violations of either rule. The lock serialises admission, not the calls themselves — once acquire() returns, callers run concurrently. Holding the lock while sleeping is deliberate: it keeps callers in FIFO order and stops a stampede of re-checks when a window opens.
Verify: record send timestamps under load and check every window against every limit; there should be no violations.
3. See why sequential acquisition violates limits¶
The tempting composition acquires limiters one after another:
async def sequential(limits):
for limit in limits:
while (w := limit.wait_time(time.monotonic())) > 0:
await asyncio.sleep(w)
limit.record(time.monotonic()) # recorded now, sent later
When the per-second limiter admits the call and records it at time t, but the per-3-second limiter then makes it wait until t + 0.8 s, the per-second log believes a call happened at t that actually happened at t + 0.8. A second later, that phantom entry expires early and lets an extra call through — which is how sequential acquisition produced 2 violations of the 5-per-second rule in 30 calls. The reverse failure also exists: a call admitted by the first limiter and then cancelled while waiting on the second has consumed capacity it never used.
Verify: run the violation check against sequential acquisition with your real limits; any non-zero result is this bug.
4. Add a concurrency cap alongside the rates¶
Many APIs also limit simultaneous requests. A concurrency cap is not a rate — it is released when the call finishes — so combine it as a separate stage, acquired after the rate limits and held for the call's duration:
class ApiGate:
def __init__(self, rates: MultiLimit, max_in_flight: int) -> None:
self.rates = rates
self.slots = asyncio.Semaphore(max_in_flight)
async def __aenter__(self):
await self.slots.acquire() # hold a slot for the whole call
try:
await self.rates.acquire() # then wait for every rate rule
except BaseException:
self.slots.release()
raise
return self
async def __aexit__(self, *exc):
self.slots.release()
gate = ApiGate(api_limits, max_in_flight=4)
async def call_api(session, url):
async with gate:
async with session.get(url) as r:
return await r.json()
Acquiring the slot first means a caller only consumes rate capacity when it can actually send, avoiding the "admitted by the rate, then stuck waiting for a slot" skew from step 3. The semaphore pattern itself is covered in limiting concurrent requests with asyncio.Semaphore.
Verify: under load, in-flight calls never exceed the cap and send timestamps satisfy every rate rule.
5. Coordinate limits across processes¶
When several workers share one API key, each process's limiter only sees its own calls. Move the check-and-record into Redis as a single atomic script over all windows:
MULTI = """
local now = tonumber(ARGV[1])
for i = 1, #KEYS do
local limit, window = tonumber(ARGV[i*2]), tonumber(ARGV[i*2+1])
redis.call('ZREMRANGEBYSCORE', KEYS[i], '-inf', now - window)
if redis.call('ZCARD', KEYS[i]) >= limit then
local oldest = redis.call('ZRANGE', KEYS[i], 0, 0, 'WITHSCORES')[2]
return tostring(tonumber(oldest) + window - now) -- wait required
end
end
for i = 1, #KEYS do
redis.call('ZADD', KEYS[i], now, ARGV[#ARGV])
redis.call('EXPIRE', KEYS[i], math.ceil(tonumber(ARGV[i*2+1])))
end
return '0'
"""
async def acquire_global(r, member: str) -> None:
rules = [("rl:sec", 10, 1), ("rl:min", 500, 60)]
keys = [k for k, _, _ in rules]
while True:
args = [time.time()] + [x for _, lim, win in rules for x in (lim, win)] + [member]
wait = float(await r.eval(MULTI, len(keys), *keys, *args))
if wait <= 0:
return
await asyncio.sleep(wait)
Lua scripts run atomically in Redis, so "check every rule, then record in every rule" holds across the whole fleet. Use a unique member per call (a UUID) so two calls at the same timestamp both count. This is the multi-window version of the single-limit script in the sliding-window guide linked above.
Verify: with four workers sharing the key, the API's own rate-limit headers never report an exceeded limit.
Verification¶
Multi-limit enforcement is correct when:
- Every rule is checked before any rule records, and all record the same timestamp.
- Measured send times satisfy every window under load, with zero violations.
- Concurrency caps are a separate stage acquired before the rate rules.
- Shared keys use an atomic shared check across processes.
Diagnostic Hook: export, per rule, the current usage as a fraction of its limit and the time callers spend waiting on it. The rule with the highest usage is the binding constraint; if the per-day rule binds by mid-afternoon, no amount of per-second tuning will help, and the conversation with the provider should be about the daily quota.
Pitfalls & edge cases¶
- Acquiring limiters one after another. Earlier rules record too early; measured violations follow.
- Cancellation between stages. Capacity consumed by a cancelled caller is not reclaimed in sequential designs.
- Per-process limiters for a shared key. The provider sees the sum.
- Clock differences in distributed limits. Use Redis server time, or keep hosts synchronised.
Frequently Asked Questions¶
How do I enforce a per-second and a per-minute rate limit together?
Ask every limit how long the caller must wait, sleep for the longest, re-check, and when every limit allows the call, record it in all of them with the same timestamp. In testing this produced no violations, while acquiring limiters one after another produced two.
Why does chaining rate limiters cause violations?
The first limiter records the call when it admits it, but the call is only sent after the later limiters allow it. The early timestamp expires too soon and lets an extra call through later.
How do I combine a rate limit with a concurrency limit?
Acquire the concurrency slot first and hold it for the call, then wait for the rate rules, so rate capacity is only consumed by callers that can actually send.
How do I share multiple rate limits across processes?
Run the check-all-then-record-all logic as a single Lua script in Redis over sorted sets, one per rule. Scripts execute atomically, so the rules hold across every process.
Related¶
- Rate Limiting & Throttling — up to the topic overview.
- Handling 429 Retry-After responses in async clients — what to do when a limit is hit anyway.
- Concurrent Execution & Worker Patterns — the section overview.