Retrying LLM API Calls on 429 and 529 Errors¶
LLM APIs fail transiently more often than most HTTP services: 429 when you exceed a rate limit, 529 overloaded_error when the provider is short of capacity, occasional 5xxs and connection resets. The official SDKs retry some of these for you, and knowing exactly which — and which they cannot — decides what your own code must handle. Measured with the anthropic SDK 1.11.0 on Python 3.14 against a local mock of the Messages API: with 30% of requests answered by 529, max_retries=0 succeeded 70% of the time (209 of 300), the default max_retries=2 succeeded 97.3% (292), and max_retries=5 succeeded 99% (297) but took 13.7 s for the batch instead of 1.7 s. Against an endpoint that always returned 529 with retry-after: 2, the SDK's three attempts came at 0, 2.0 and 4.01 s — exactly as instructed. Without retry-after, the gaps between six attempts were about 0.44, 0.9, 1.8, 3.6 and 7.4 s. The SDK retried 408, 409, 429 and every 5xx. It did not retry an overloaded_error that arrived as an event after streaming had begun, nor a timeout that fired mid-stream. This guide fills those gaps without creating retry storms.
Prerequisites¶
- Python 3.11+,
pip install anthropic. - Backoff and jitter, from exponential backoff with jitter in asyncio.
- The topic overview, Concurrent LLM API Calls.
1. Know what the SDK already retries¶
AsyncAnthropic retries a request up to max_retries times (default 2) when the response status is 408, 409, 429 or 500 and above, or when the connection fails before a response arrives. A response header x-should-retry: true or false overrides the status rule. The delay honours retry-after (and retry-after-ms) when the server sends it; otherwise it backs off exponentially from 0.5 s, capped at 8 s, with jitter:
from anthropic import AsyncAnthropic
client = AsyncAnthropic(max_retries=2) # the default
batch_client = client.with_options(max_retries=5) # a copy for work that can wait
Measured on the always-529 endpoint with max_retries=5: six attempts, gaps of 0.43–0.45 s, 0.84–1.0 s, 1.78–1.87 s, 3.47–3.96 s and 7.15–7.79 s, about 14 s in total before OverloadedError was raised. With retry-after: 2 the gaps were exactly 2.0 s. Everything the SDK retries is invisible to your code except as latency; everything it gives up on arrives as a subclass of anthropic.APIStatusError — RateLimitError for 429, OverloadedError for 529, InternalServerError for other 5xx.
Verify: log exc.status_code and the elapsed time for every APIStatusError you catch; the elapsed time shows how much retrying already happened.
2. Choose retry counts by caller, not globally¶
Interactive and batch callers want different trade-offs. A chat user waiting for a reply would rather see an error after a few seconds than a spinner for fourteen; a nightly job would rather wait than fail:
interactive = AsyncAnthropic(max_retries=1, timeout=30.0)
batch = AsyncAnthropic(max_retries=6)
# or per call, without a second client
msg = await client.with_options(max_retries=0).messages.create(...)
With independent failures, each retry multiplies the failure probability by the failure rate: at 30% overload, no retries left 30% failing, two left about 3% (measured 2.7%), five about 0.2% (measured 1%). The cost is latency on the failing path: two retries without retry-after add up to about 1.4 s, five add about 14 s. Interactive paths should also wrap the whole call in a deadline, as in setting per-attempt and total timeouts for retries, so retries cannot outlast the user's patience.
Verify: interactive endpoints fail within their latency budget under a simulated overload, and batch jobs complete.
3. Retry mid-stream failures yourself¶
Once a stream has started, the HTTP status is 200 and the SDK's retry logic is behind you. An error event or a dropped connection mid-stream ends the call with an exception and partial output:
async def stream_with_retry(prompt: str, attempts: int = 3) -> str:
for attempt in range(attempts):
parts: list[str] = []
try:
async with client.messages.stream(model=MODEL, max_tokens=1024,
messages=[{"role": "user", "content": prompt}]) as s:
async for text in s.text_stream:
parts.append(text)
return "".join(parts)
except anthropic.APIStatusError as exc:
if exc.status_code not in (200, 429, 500, 502, 503, 529) or attempt == attempts - 1:
raise
except httpx2.TransportError:
if attempt == attempts - 1:
raise
await asyncio.sleep(min(8.0, 0.5 * 2 ** attempt) * random.uniform(0.5, 1.0))
raise AssertionError("unreachable")
Measured: an overloaded_error event after 20 tokens raised APIStatusError with status_code 200 after 0.10 s, and the mock recorded exactly one request — no retry. The 200 in the retryable set is deliberate: it is how mid-stream errors present. Retrying restarts generation from the beginning, so discard the partial text (or, for a user-facing stream, tell the client to clear it). Whether to retry also depends on whether anything was already shown to the user; for relays, see relaying LLM token streams through FastAPI.
Verify: a mock that fails mid-stream once and then succeeds produces one complete answer, not a concatenation of partial ones.
4. Coordinate retries across concurrent callers¶
Per-call retries are independent: when a rate limit is hit, every concurrent caller discovers it separately and schedules its own retry. With many callers that becomes a synchronized wave — measured in the fan-out guide as 480 rejected attempts for 200 calls with unbounded concurrency. Two measures keep retries from amplifying load:
class RetryBudget:
"""Allow retries only up to a fraction of recent first attempts."""
def __init__(self, ratio: float = 0.2) -> None:
self.ratio, self.attempts, self.retries = ratio, 0, 0
def first_attempt(self) -> None:
self.attempts += 1
def can_retry(self) -> bool:
if self.retries < self.ratio * max(self.attempts, 10):
self.retries += 1
return True
return False
A bounded concurrency limit stops the wave at its source, and a retry budget caps the extra load retries can add during a provider-side incident — when 529s are widespread, retrying everything multiplies traffic exactly when capacity is shortest. Use the SDK's retries for the routine case (max_retries small) and a budget in your own retry loop for the rest; the general pattern is in implementing retry budgets to prevent retry storms.
Verify: during a simulated provider outage, total request rate stays within the budget's multiple of the normal rate.
5. Do not retry what cannot succeed¶
Some errors are permanent for a given request, and retrying them only wastes quota and time:
NON_RETRYABLE = {400, 401, 403, 404, 413, 422}
def should_retry(exc: Exception) -> bool:
if isinstance(exc, anthropic.APIStatusError):
return exc.status_code not in NON_RETRYABLE
return isinstance(exc, (anthropic.APIConnectionError, httpx2.TransportError, TimeoutError))
A 400 for a malformed request or a prompt over the context window, 401/403 for credentials, 413 for an oversized payload — none of these will change on retry. The SDK already refrains from retrying them; your own loops, especially those wrapping streams, should apply the same rule. Log them with enough of the request to reproduce: they are bugs or configuration problems, not weather. The broader taxonomy is in classifying retryable errors in async clients.
Verify: an invalid request fails once, immediately, and is logged as an error rather than retried.
Verification¶
LLM retries are correct when:
- SDK retries are configured per caller type, small for interactive paths and larger for batch.
- Mid-stream failures are retried by application code, which treats status 200 as retryable and discards partial output.
- Permanent errors are never retried, and are logged as failures.
- Retries are bounded across callers by concurrency limits and a retry budget.
Diagnostic Hook: export counts of RateLimitError, OverloadedError and mid-stream APIStatusError with status 200, plus the elapsed time of failed calls. Elapsed times near the SDK's full backoff schedule mean retries are being exhausted — a capacity or rate-limit problem, not a transient one.
Pitfalls & edge cases¶
- Assuming the SDK retries streams. Measured: a mid-stream error raised immediately, one request, no retry.
- Large
max_retrieson interactive paths. Five retries withoutretry-aftertook about 14 s. - Ignoring status 200 errors. Mid-stream failures carry the original response's status.
- Retrying on every caller during an outage. It multiplies load when capacity is scarcest.
Frequently Asked Questions¶
Does the Anthropic Python SDK retry 429 and 529 errors?
Yes. It retries 408, 409, 429 and 5xx responses up to max_retries times (default 2), honouring retry-after when present and otherwise backing off exponentially from about 0.5 s to a cap of 8 s. In testing it waited exactly 2.0 s between attempts for retry-after: 2.
How many retries should I use for LLM API calls?
Two recovered most transient failures in testing — 97.3% success at a 30% overload rate, against 70% with none. Use fewer for interactive requests with a deadline and more for batch jobs.
Why wasn't my streaming request retried?
Errors after the stream has started arrive as events on a 200 response, so the SDK raises APIStatusError with status_code 200 and does not retry. Catch it, discard partial output and retry in your own code.
Which LLM API errors should never be retried?
400 (invalid request or over the context window), 401 and 403 (credentials), 404 and 413 (oversized payload): they fail the same way every time.
Related¶
- Concurrent LLM API Calls — up to the topic overview.
- Streaming LLM tokens asynchronously — where mid-stream errors come from.
- Concurrent Execution & Worker Patterns — the section overview.