Rate Limiting LLM Calls by Tokens per Minute¶
LLM providers limit more than requests per minute. They limit tokens — input and output — per minute, so ten small requests and ten enormous ones consume very different shares of the same budget, and a limiter that counts requests or caps concurrency cannot see the difference. The result is a stream of 429 rate_limit_error responses even though the client looks well-behaved. To measure it, 60 calls each reserving about 207 tokens (a short prompt plus max_tokens=200) were sent through the anthropic SDK on Python 3.14 to a local mock of the Messages API that enforced 4,000 tokens per sliding 10-second window — a scaled-down tokens-per-minute limit, so each run took seconds rather than minutes. A Semaphore(10) alone completed 36 calls and failed 24 after the SDK's retries, with 82 429 responses in 118 attempts. A token bucket that started full and refilled at the window's average rate still produced 60 429s and 12 failures, because its burst plus its refill exceeded what the window allowed. A bucket sized so burst plus one window of refill equalled the limit, and a client-side sliding window identical to the server's, both completed all 60 with zero 429s, in 38.3 s and 30.5 s. This guide builds a token-aware limiter that matches the budget it is protecting.
Prerequisites¶
- Python 3.11+,
pip install anthropic. - Token buckets in asyncio, from token bucket rate limiter for asyncio clients.
- The topic overview, Concurrent LLM API Calls.
1. See why request and concurrency limits miss¶
A semaphore bounds how many calls are in flight, not how many tokens they consume:
sem = asyncio.Semaphore(10)
async def summarize(doc: str) -> str:
async with sem:
msg = await client.messages.create(model="claude-sonnet-5-5", max_tokens=200,
messages=[{"role": "user", "content": doc}])
return msg.content[0].text
Measured against the token budget: 36 of 60 calls succeeded, 24 failed with 429 after the SDK's two retries each, and the client made 118 HTTP attempts — most of them doomed. Each call finished in well under a second, so ten concurrent slots pushed far more than 4,000 tokens into a 10-second window. Retries made it worse: every rejected attempt was immediately replaced by another. A semaphore is still useful — the provider also limits concurrent requests, and fanning out with bounded concurrency covers that — but it is not a token limiter.
Verify: compare your 429 rate with your semaphore size; if 429s persist at any concurrency, the binding limit is tokens.
2. Reserve tokens before each call¶
A token limiter needs a number for each call before it is sent. Use an estimate of the input plus the full max_tokens, because the output length is unknown until the response ends:
import json
def estimate_input_tokens(messages: list[dict]) -> int:
# Rough: ~4 characters per token for English text. Use the provider's
# token-counting endpoint when estimates must be exact.
return max(1, len(json.dumps(messages)) // 4)
async def create_limited(limiter, **kwargs):
reserve = estimate_input_tokens(kwargs["messages"]) + kwargs["max_tokens"]
await limiter.acquire(reserve)
return await client.messages.create(**kwargs)
Reserving max_tokens is conservative: a call that asks for 2,000 tokens and uses 150 holds budget it does not need. That is the safe direction — under-reserving produces 429s — and it rewards setting max_tokens close to what each task actually needs rather than a uniform large value. The Anthropic API offers a token-counting endpoint for exact input counts; the character heuristic is adequate for budgeting as long as it errs high.
Verify: the sum of reservations per window, logged by the limiter, stays at or below the provider's limit.
3. Match the limiter to the server's window¶
The second row of the table is the instructive failure. A token bucket that starts full at capacity C and refills at rate r can spend C + r × W in a window of length W. With C = 4,000 and r = 400/s, that is 8,000 tokens in the first ten seconds against a server that allows 4,000 — so it produced 60 429s even though its long-run average matched the limit exactly. Two shapes avoided that:
class TokenBucket:
def __init__(self, capacity: float, per_second: float) -> None:
self.capacity = self.tokens = capacity
self.rate, self.updated = per_second, time.monotonic()
self._lock = asyncio.Lock() # FIFO: large requests are not starved
async def acquire(self, n: float) -> None:
async with self._lock:
while True:
now = time.monotonic()
self.tokens = min(self.capacity, self.tokens + (now - self.updated) * self.rate)
self.updated = now
if self.tokens >= n:
self.tokens -= n
return
await asyncio.sleep((n - self.tokens) / self.rate)
LIMIT, WINDOW = 4000, 10.0
bucket = TokenBucket(capacity=0.2 * LIMIT, per_second=0.8 * LIMIT / WINDOW) # C + r*W == LIMIT
The sized bucket completed all 60 calls with no 429s in 38.3 s. A sliding window that tracks reservations over the last W seconds — the same algorithm the mock used — did the same in 30.5 s, because it can use the whole budget in a burst when the window is empty, exactly as the server permits:
class SlidingWindow:
def __init__(self, limit: int, window: float) -> None:
self.limit, self.window = limit, window
self.spent: collections.deque[tuple[float, int]] = collections.deque()
self._lock = asyncio.Lock()
async def acquire(self, n: int) -> None:
async with self._lock:
while True:
now = time.monotonic()
while self.spent and now - self.spent[0][0] >= self.window:
self.spent.popleft()
if sum(t for _, t in self.spent) + n <= self.limit:
self.spent.append((now, n))
return
await asyncio.sleep(self.spent[0][0] + self.window - now + 0.01)
Matching the window exactly is not quite enough under heavier load. In a later run with 20 concurrent streams, retries and a 20,000-token window, an identical client window still drew 29 token 429s out of 386 attempts: each reservation was timestamped on the client a moment before the server recorded the request, so it aged out of the client's window first. Reserving against 95% of the limit over the window plus 0.25 s brought that to zero. Providers document their algorithms to varying degrees; the safe rule is that whatever the client allows in any window must not exceed what the server allows in that window. The sized bucket guarantees that for any server algorithm with the same average rate, at some cost in throughput.
Verify: run a load test at full budget for three windows; the 429 count is zero.
4. Share one limiter across every caller¶
A limiter protects a budget only if every call that spends the budget goes through it. In a service that means one instance per API key per process — and a smaller share of the budget per process when several processes share a key:
class LLMGateway:
def __init__(self, client: AsyncAnthropic, tokens_per_minute: int, processes: int,
max_concurrency: int) -> None:
share = tokens_per_minute / processes
self.client = client
self.tokens = SlidingWindow(int(share), 60.0)
self.slots = asyncio.Semaphore(max_concurrency)
async def create(self, **kwargs):
await self.tokens.acquire(estimate_input_tokens(kwargs["messages"]) + kwargs["max_tokens"])
async with self.slots:
return await self.client.messages.create(**kwargs)
Acquiring tokens before the concurrency slot keeps a call that is waiting for budget from occupying a slot that a cheaper call could use. Splitting the budget statically between processes is simple and leaves some capacity unused when load is uneven; a shared limiter in Redis, as in sliding window rate limiting with Redis and asyncio, uses all of it at the cost of a round trip per call.
Verify: every code path that calls the LLM goes through the gateway; grep finds no direct messages.create elsewhere.
5. Keep 429 handling as the backstop¶
Even a correct limiter will see occasional 429s — a clock skew, another client on the same key, a provider-side adjustment. Let the SDK retry them, and feed what the server says back into the limiter:
class SlidingWindow:
...
async def pause(self, seconds: float) -> None:
async with self._lock: # acquirers queue behind the lock meanwhile
await asyncio.sleep(seconds)
async def create(self, **kwargs):
reserve = estimate_input_tokens(kwargs["messages"]) + kwargs["max_tokens"]
await self.tokens.acquire(reserve)
try:
async with self.slots:
return await self.client.messages.create(**kwargs)
except anthropic.RateLimitError as exc:
RATE_LIMITED.inc()
retry_after = float(exc.response.headers.get("retry-after", "1"))
await self.tokens.pause(retry_after) # stop everyone, not just this caller
raise
A pause that blocks all acquirers for the server's retry-after turns one 429 into a coordinated slowdown instead of every concurrent caller discovering the limit separately. The SDK's own retries honour retry-after — measured in the retry guide — but they act per call, not across callers.
Verify: after an injected 429, all callers wait the indicated time, and the next window shows no further 429s.
Verification¶
Token-based limiting is working when:
- Each call reserves estimated input tokens plus
max_tokensbefore it is sent. - The limiter's worst-case window spend is at most the provider's limit — a matching sliding window, or a bucket with
C + r × W ≤ limit. - One limiter per key per process gates every call, with the budget divided between processes.
- 429s are rare, counted, and pause all callers for the server's
retry-after.
Diagnostic Hook: log reserved tokens per window from the limiter and actual tokens per window from response usage. If 429s appear while reserved tokens are under the limit, another client shares the key or the estimate is low; if actual usage is far below reservations, lower max_tokens to recover throughput.
Pitfalls & edge cases¶
- Concurrency caps as rate limits. Measured: 24 of 60 calls failed with a semaphore alone.
- Full-burst token buckets. Measured: 60 429s despite the correct average rate.
- Reserving only input tokens. Output tokens count against the same budget.
- Per-call retries without coordination. Every caller rediscovers the limit separately.
Frequently Asked Questions¶
How do I rate limit LLM API calls by tokens per minute in Python?
Before each call, reserve estimated input tokens plus max_tokens from a limiter whose worst-case spend in any window is at most the provider's limit. In testing a client sliding window matching the server completed 60 of 60 calls with no 429s.
Why do I get 429 errors with a concurrency limit on LLM calls?
A semaphore limits calls in flight, not tokens per minute. Fast calls cycle through the slots and exceed the token budget: a Semaphore(10) alone failed 24 of 60 calls in testing.
Why does my token bucket still get rate limited?
A bucket that starts full can spend its whole capacity plus a window of refill in the first window, double the limit when both equal the budget. Size it so capacity plus rate times window equals the limit.
Should I reserve max_tokens or actual output tokens?
Reserve max_tokens, since output length is unknown until the call ends; it errs on the safe side. Set max_tokens close to what each task needs so reservations are not wasted.
Related¶
- Concurrent LLM API Calls — up to the topic overview.
- Fanning out LLM calls with bounded concurrency — the concurrent-request side of the limit.
- Concurrent Execution & Worker Patterns — the section overview.