Skip to content

Retrying LLM API Calls on 429 and 529 Errors

LLM APIs fail transiently more often than most HTTP services: 429 when you exceed a rate limit, 529 overloaded_error when the provider is short of capacity, occasional 5xxs and connection resets. The official SDKs retry some of these for you, and knowing exactly which — and which they cannot — decides what your own code must handle. Measured with the anthropic SDK 1.11.0 on Python 3.14 against a local mock of the Messages API: with 30% of requests answered by 529, max_retries=0 succeeded 70% of the time (209 of 300), the default max_retries=2 succeeded 97.3% (292), and max_retries=5 succeeded 99% (297) but took 13.7 s for the batch instead of 1.7 s. Against an endpoint that always returned 529 with retry-after: 2, the SDK's three attempts came at 0, 2.0 and 4.01 s — exactly as instructed. Without retry-after, the gaps between six attempts were about 0.44, 0.9, 1.8, 3.6 and 7.4 s. The SDK retried 408, 409, 429 and every 5xx. It did not retry an overloaded_error that arrived as an event after streaming had begun, nor a timeout that fired mid-stream. This guide fills those gaps without creating retry storms.

Prerequisites

1. Know what the SDK already retries

AsyncAnthropic retries a request up to max_retries times (default 2) when the response status is 408, 409, 429 or 500 and above, or when the connection fails before a response arrives. A response header x-should-retry: true or false overrides the status rule. The delay honours retry-after (and retry-after-ms) when the server sends it; otherwise it backs off exponentially from 0.5 s, capped at 8 s, with jitter:

from anthropic import AsyncAnthropic

client = AsyncAnthropic(max_retries=2)            # the default
batch_client = client.with_options(max_retries=5) # a copy for work that can wait

Measured on the always-529 endpoint with max_retries=5: six attempts, gaps of 0.43–0.45 s, 0.84–1.0 s, 1.78–1.87 s, 3.47–3.96 s and 7.15–7.79 s, about 14 s in total before OverloadedError was raised. With retry-after: 2 the gaps were exactly 2.0 s. Everything the SDK retries is invisible to your code except as latency; everything it gives up on arrives as a subclass of anthropic.APIStatusError — RateLimitError for 429, OverloadedError for 529, InternalServerError for other 5xx.

Verify: log exc.status_code and the elapsed time for every APIStatusError you catch; the elapsed time shows how much retrying already happened.

Success rate with 30% of requests overloaded (529) 3 horizontal bars comparing max_retries=0 with the others. Success rate with 30% of requests overloaded (529) max_retries=0 70% (209/300), 0.4 s max_retries=2 (default) 97.3% (292/300), 1.7 s max_retries=5 99% (297/300), 13.7 s Local mock; anthropic 1.11.0. Each retry is an independent 30% chance of failing again. Two retries recover most transient failures; more retries buy little and cost latency.

2. Choose retry counts by caller, not globally

Interactive and batch callers want different trade-offs. A chat user waiting for a reply would rather see an error after a few seconds than a spinner for fourteen; a nightly job would rather wait than fail:

interactive = AsyncAnthropic(max_retries=1, timeout=30.0)
batch = AsyncAnthropic(max_retries=6)

# or per call, without a second client
msg = await client.with_options(max_retries=0).messages.create(...)

With independent failures, each retry multiplies the failure probability by the failure rate: at 30% overload, no retries left 30% failing, two left about 3% (measured 2.7%), five about 0.2% (measured 1%). The cost is latency on the failing path: two retries without retry-after add up to about 1.4 s, five add about 14 s. Interactive paths should also wrap the whole call in a deadline, as in setting per-attempt and total timeouts for retries, so retries cannot outlast the user's patience.

Verify: interactive endpoints fail within their latency budget under a simulated overload, and batch jobs complete.

3. Retry mid-stream failures yourself

Once a stream has started, the HTTP status is 200 and the SDK's retry logic is behind you. An error event or a dropped connection mid-stream ends the call with an exception and partial output:

async def stream_with_retry(prompt: str, attempts: int = 3) -> str:
    for attempt in range(attempts):
        parts: list[str] = []
        try:
            async with client.messages.stream(model=MODEL, max_tokens=1024,
                                              messages=[{"role": "user", "content": prompt}]) as s:
                async for text in s.text_stream:
                    parts.append(text)
            return "".join(parts)
        except anthropic.APIStatusError as exc:
            if exc.status_code not in (200, 429, 500, 502, 503, 529) or attempt == attempts - 1:
                raise
        except httpx2.TransportError:
            if attempt == attempts - 1:
                raise
        await asyncio.sleep(min(8.0, 0.5 * 2 ** attempt) * random.uniform(0.5, 1.0))
    raise AssertionError("unreachable")

Measured: an overloaded_error event after 20 tokens raised APIStatusError with status_code 200 after 0.10 s, and the mock recorded exactly one request — no retry. The 200 in the retryable set is deliberate: it is how mid-stream errors present. Retrying restarts generation from the beginning, so discard the partial text (or, for a user-facing stream, tell the client to clear it). Whether to retry also depends on whether anything was already shown to the user; for relays, see relaying LLM token streams through FastAPI.

Verify: a mock that fails mid-stream once and then succeeds produces one complete answer, not a concatenation of partial ones.

Who retries each failure A grid of 5 rows by 3 columns. Who retries each failure failure raised as retried by the SDK 429 / 529 / 5xx before the stream RateLimitError, OverloadedError, ... yes, honours retry-after connection failed before a response APIConnectionError yes timeout, non-streaming create() APITimeoutError (after 10.2 s with timeout=3) yes error event after streaming began APIStatusError, status_code 200 no read timeout mid-stream httpx2.ReadTimeout no anthropic 1.11.0 against a local mock; the last two rows are your code's responsibility.

4. Coordinate retries across concurrent callers

Per-call retries are independent: when a rate limit is hit, every concurrent caller discovers it separately and schedules its own retry. With many callers that becomes a synchronized wave — measured in the fan-out guide as 480 rejected attempts for 200 calls with unbounded concurrency. Two measures keep retries from amplifying load:

class RetryBudget:
    """Allow retries only up to a fraction of recent first attempts."""
    def __init__(self, ratio: float = 0.2) -> None:
        self.ratio, self.attempts, self.retries = ratio, 0, 0

    def first_attempt(self) -> None:
        self.attempts += 1

    def can_retry(self) -> bool:
        if self.retries < self.ratio * max(self.attempts, 10):
            self.retries += 1
            return True
        return False

A bounded concurrency limit stops the wave at its source, and a retry budget caps the extra load retries can add during a provider-side incident — when 529s are widespread, retrying everything multiplies traffic exactly when capacity is shortest. Use the SDK's retries for the routine case (max_retries small) and a budget in your own retry loop for the rest; the general pattern is in implementing retry budgets to prevent retry storms.

Verify: during a simulated provider outage, total request rate stays within the budget's multiple of the normal rate.

5. Do not retry what cannot succeed

Some errors are permanent for a given request, and retrying them only wastes quota and time:

NON_RETRYABLE = {400, 401, 403, 404, 413, 422}


def should_retry(exc: Exception) -> bool:
    if isinstance(exc, anthropic.APIStatusError):
        return exc.status_code not in NON_RETRYABLE
    return isinstance(exc, (anthropic.APIConnectionError, httpx2.TransportError, TimeoutError))

A 400 for a malformed request or a prompt over the context window, 401/403 for credentials, 413 for an oversized payload — none of these will change on retry. The SDK already refrains from retrying them; your own loops, especially those wrapping streams, should apply the same rule. Log them with enough of the request to reproduce: they are bugs or configuration problems, not weather. The broader taxonomy is in classifying retryable errors in async clients.

Verify: an invalid request fails once, immediately, and is logged as an error rather than retried.

Should this LLM failure be retried, and by whom? A decision on When and how did it fail with 4 outcomes. Should this LLM failure be retried, and by whom? When and how did it fail? 429 / 529 / 5xx before streaming SDK retries max_retries 1 interactive, ~6 batch after the stream started your loop retries discard partial output 400, 401, 403, 413 never retry fix the request many callers failing at once retry budget cap extra load The SDK covers the common case; streams and storms are yours.

Verification

LLM retries are correct when:

  • SDK retries are configured per caller type, small for interactive paths and larger for batch.
  • Mid-stream failures are retried by application code, which treats status 200 as retryable and discards partial output.
  • Permanent errors are never retried, and are logged as failures.
  • Retries are bounded across callers by concurrency limits and a retry budget.

Diagnostic Hook: export counts of RateLimitError, OverloadedError and mid-stream APIStatusError with status 200, plus the elapsed time of failed calls. Elapsed times near the SDK's full backoff schedule mean retries are being exhausted — a capacity or rate-limit problem, not a transient one.

Pitfalls & edge cases

  • Assuming the SDK retries streams. Measured: a mid-stream error raised immediately, one request, no retry.
  • Large max_retries on interactive paths. Five retries without retry-after took about 14 s.
  • Ignoring status 200 errors. Mid-stream failures carry the original response's status.
  • Retrying on every caller during an outage. It multiplies load when capacity is scarcest.

Frequently Asked Questions

Does the Anthropic Python SDK retry 429 and 529 errors?

Yes. It retries 408, 409, 429 and 5xx responses up to max_retries times (default 2), honouring retry-after when present and otherwise backing off exponentially from about 0.5 s to a cap of 8 s. In testing it waited exactly 2.0 s between attempts for retry-after: 2.

How many retries should I use for LLM API calls?

Two recovered most transient failures in testing — 97.3% success at a 30% overload rate, against 70% with none. Use fewer for interactive requests with a deadline and more for batch jobs.

Why wasn't my streaming request retried?

Errors after the stream has started arrive as events on a 200 response, so the SDK raises APIStatusError with status_code 200 and does not retry. Catch it, discard partial output and retry in your own code.

Which LLM API errors should never be retried?

400 (invalid request or over the context window), 401 and 403 (credentials), 404 and 413 (oversized payload): they fail the same way every time.