Retrying Async Calls with tenacity¶
tenacity is the most widely used retry library in Python, and its decorators work on coroutines unchanged: @retry on an async def retries with asyncio.sleep between attempts. Its defaults and some of its stop conditions do not mean what a quick read suggests. Tested with tenacity 9.1.4: a bare @retry(stop=stop_after_attempt(3)) retried a ValueError — a bug in the caller's input — just as readily as a connection error; after the last attempt it raised RetryError, not the original exception, unless reraise=True was set; and stop_after_delay(1.0) with 0.3 s attempts let the sequence run 1.30 s, because it only checks after an attempt finishes. Wrapping the loop in asyncio.timeout(1.0) stopped it at 1.00 s. Cancellation behaved well: a cancel during backoff ended the call immediately, and a CancelledError raised by the wrapped call was never retried. This guide configures tenacity for async code so it retries only what it should, for only as long as it should.
Prerequisites¶
- Python 3.11+,
pip install tenacity(tested with 9.1.4). - Which errors to retry, from classifying retryable errors in async clients.
- Backoff, from exponential backoff with jitter in asyncio.
1. Restrict retries to transient errors¶
Without a retry= condition, tenacity retries every Exception. Name the errors that are worth another attempt:
import httpx
from tenacity import (retry, retry_if_exception, stop_after_attempt,
wait_exponential_jitter)
TRANSIENT = (httpx.ConnectError, httpx.ReadTimeout, httpx.RemoteProtocolError)
def is_transient(exc: BaseException) -> bool:
if isinstance(exc, httpx.HTTPStatusError):
return exc.response.status_code in (429, 502, 503, 504)
return isinstance(exc, TRANSIENT)
@retry(
retry=retry_if_exception(is_transient),
stop=stop_after_attempt(4),
wait=wait_exponential_jitter(initial=0.1, max=2.0, jitter=0.1),
reraise=True,
)
async def get_profile(client: httpx.AsyncClient, user_id: int) -> dict:
response = await client.get(f"/users/{user_id}")
response.raise_for_status()
return response.json()
Tested: a bare @retry retried a ValueError three times. Retrying bugs and bad input multiplies load and delays the error the caller needs to see. A predicate function (retry_if_exception) can inspect status codes and error attributes that a type list cannot express.
The predicate is also the right place to honour server hints. A 429 or 503 with a Retry-After header tells the client when to come back; tenacity's wait can read it from the failed attempt's exception (retry_state.outcome.exception()) and wait that long instead of the computed backoff, as covered in handling 429 Retry-After responses in async clients.
Verify: a unit test that raises a 400 or ValueError from the wrapped call sees exactly one attempt.
2. Re-raise the original exception¶
By default, exhausted retries raise tenacity.RetryError, which wraps the last attempt. Callers that catch httpx.HTTPError or a domain exception miss it:
try:
await get_profile(client, 42)
except httpx.HTTPError:
... # without reraise=True this never runs: RetryError is raised
@retry(..., reraise=True) # tested: the original ConnectionError propagates
async def get_profile(client: httpx.AsyncClient, user_id: int) -> dict: ...
Tested: without reraise=True, the caller received RetryError with the ConnectionError inside; with it, the ConnectionError itself. reraise=True keeps retries an implementation detail of the function — callers handle the same exceptions whether or not it retried. If you need the attempt count in the error, add it as a note in a before_sleep or retry_error_callback hook rather than changing the exception type.
Verify: exceptions reaching callers after exhausted retries are the same types the unwrapped call raises.
3. Bound total time with asyncio.timeout, not stop_after_delay¶
stop_after_delay is checked after each failed attempt, so the last attempt can run past the limit by its full duration. Measured: stop_after_delay(1.0) with attempts that timed out at 0.3 s ran for 1.30 s. An asyncio.timeout around the whole retry loop is a real deadline:
from tenacity import AsyncRetrying, retry_if_exception_type, wait_exponential_jitter
async def fetch_with_budget(client, url: str, budget: float = 1.0, per_attempt: float = 0.3):
async with asyncio.timeout(budget): # hard deadline for everything
async for attempt in AsyncRetrying(
retry=retry_if_exception_type((TimeoutError, httpx.TransportError)),
wait=wait_exponential_jitter(initial=0.05, max=0.2),
reraise=True,
):
with attempt:
async with asyncio.timeout(per_attempt): # bound each attempt
return await client.get(url)
Measured: the same three attempts ended at exactly 1.00 s. tenacity 8.3+ also offers stop_before_delay, which refuses to start an attempt that could not finish in time, but it cannot cut short an attempt in progress; the outer timeout can. The budget arithmetic for per-attempt and total timeouts is in setting per-attempt and total timeouts for retries.
Verify: with a dependency that never answers, the call fails at the budget, not at the budget plus one attempt.
4. Let cancellation through¶
Retry logic must never turn a cancellation into another attempt. tenacity handles this correctly for coroutines:
task = asyncio.create_task(get_profile(client, 42))
await asyncio.sleep(0.2)
task.cancel() # tested: CancelledError immediately, even mid-backoff
Tested: cancelling during a long backoff sleep raised CancelledError at once, after one attempt; a wrapped call that itself raised CancelledError was not retried, because tenacity's retry conditions match Exception, and CancelledError is a BaseException. Keep it that way in your own predicates — never write retry_if_exception_type(BaseException) — and do not catch CancelledError inside the retried function. The same applies to the asyncio.timeout cancellation inside an attempt, which arrives as TimeoutError and is retryable by choice.
Verify: a test cancels a call during its backoff and asserts it finishes promptly with CancelledError.
5. Observe retries¶
Silent retries hide degraded dependencies until they fail outright. Log and count every retry:
import logging
from tenacity import before_sleep_log
log = logging.getLogger(__name__)
def count_retry(retry_state) -> None:
RETRIES.labels(fn=retry_state.fn.__name__,
error=type(retry_state.outcome.exception()).__name__).inc()
@retry(
retry=retry_if_exception(is_transient),
stop=stop_after_attempt(4),
wait=wait_exponential_jitter(initial=0.1, max=2.0),
before_sleep=lambda rs: (count_retry(rs), before_sleep_log(log, logging.WARNING)(rs)),
reraise=True,
)
async def get_profile(client: httpx.AsyncClient, user_id: int) -> dict: ...
before_sleep runs before each backoff with the outcome of the failed attempt. A retry counter per function and error type shows a dependency degrading long before error rates rise; tested, the decorated function's statistics also records attempt numbers and total idle time for debugging. Combine retries with a breaker and a budget so a real outage does not multiply load, as in combining circuit breakers with retries.
Verify: a dashboard shows retries per function and error type, and they correlate with dependency incidents.
Verification¶
tenacity is configured safely when:
- Every decorator has a
retry=predicate limited to transient errors. reraise=Truekeeps exception types unchanged for callers.- Total time is bounded by
asyncio.timeout, with per-attempt timeouts inside. - Retries are logged and counted, and cancellation is never retried.
Diagnostic Hook: export retries and final failures per function. A rising retry rate with stable final failures means retries are absorbing a degrading dependency — investigate it now; final failures that equal attempts × calls mean the predicate is retrying errors that never succeed.
Pitfalls & edge cases¶
- No
retry=condition. Tested:ValueErrorretried three times. - Callers catching the original type. Without
reraise=Truethey getRetryError. stop_after_delayas a deadline. Tested: 1.30 s for a 1.0 s limit.- Retrying
BaseException. Cancellation would be retried.
Frequently Asked Questions¶
Does tenacity work with async functions?
Yes. @retry on an async def retries with asyncio.sleep between attempts, and AsyncRetrying supports async for attempt loops. Tested with tenacity 9.1.4 on Python 3.14.
Why does tenacity raise RetryError instead of my exception?
That is the default after the last attempt. Set reraise=True to re-raise the original exception from the final attempt.
Does tenacity retry asyncio.CancelledError?
No. Its retry conditions match Exception, and CancelledError is a BaseException; in testing, a cancelled call and a call raising CancelledError were not retried.
How do I set a total timeout for tenacity retries?
Wrap the retry loop in async with asyncio.timeout(budget). stop_after_delay only checks between attempts and let a 1.0 s limit run to 1.30 s in testing.
Related¶
- Retry & Backoff Strategies — up to the topic overview.
- Writing a retry decorator for coroutines — when a small custom decorator is enough.
- Resilience, Cancellation & Error Handling — the section overview.