Skip to content

Retrying Async Calls with tenacity

tenacity is the most widely used retry library in Python, and its decorators work on coroutines unchanged: @retry on an async def retries with asyncio.sleep between attempts. Its defaults and some of its stop conditions do not mean what a quick read suggests. Tested with tenacity 9.1.4: a bare @retry(stop=stop_after_attempt(3)) retried a ValueError — a bug in the caller's input — just as readily as a connection error; after the last attempt it raised RetryError, not the original exception, unless reraise=True was set; and stop_after_delay(1.0) with 0.3 s attempts let the sequence run 1.30 s, because it only checks after an attempt finishes. Wrapping the loop in asyncio.timeout(1.0) stopped it at 1.00 s. Cancellation behaved well: a cancel during backoff ended the call immediately, and a CancelledError raised by the wrapped call was never retried. This guide configures tenacity for async code so it retries only what it should, for only as long as it should.

Prerequisites

1. Restrict retries to transient errors

Without a retry= condition, tenacity retries every Exception. Name the errors that are worth another attempt:

import httpx
from tenacity import (retry, retry_if_exception, stop_after_attempt,
                      wait_exponential_jitter)

TRANSIENT = (httpx.ConnectError, httpx.ReadTimeout, httpx.RemoteProtocolError)


def is_transient(exc: BaseException) -> bool:
    if isinstance(exc, httpx.HTTPStatusError):
        return exc.response.status_code in (429, 502, 503, 504)
    return isinstance(exc, TRANSIENT)


@retry(
    retry=retry_if_exception(is_transient),
    stop=stop_after_attempt(4),
    wait=wait_exponential_jitter(initial=0.1, max=2.0, jitter=0.1),
    reraise=True,
)
async def get_profile(client: httpx.AsyncClient, user_id: int) -> dict:
    response = await client.get(f"/users/{user_id}")
    response.raise_for_status()
    return response.json()

Tested: a bare @retry retried a ValueError three times. Retrying bugs and bad input multiplies load and delays the error the caller needs to see. A predicate function (retry_if_exception) can inspect status codes and error attributes that a type list cannot express.

The predicate is also the right place to honour server hints. A 429 or 503 with a Retry-After header tells the client when to come back; tenacity's wait can read it from the failed attempt's exception (retry_state.outcome.exception()) and wait that long instead of the computed backoff, as covered in handling 429 Retry-After responses in async clients.

Verify: a unit test that raises a 400 or ValueError from the wrapped call sees exactly one attempt.

tenacity 9.1.4 behaviour with coroutines, tested A grid of 7 rows by 2 columns. tenacity 9.1.4 behaviour with coroutines, tested situation result @retry with no retry= condition, ValueError retried (3 attempts) final failure, default RetryError wrapping the last exception final failure, reraise=True the original ConnectionError stop_after_delay(1.0), 0.3 s attempts ran 1.30 s asyncio.timeout(1.0) around the loop stopped at 1.00 s task cancelled during backoff CancelledError immediately call raises CancelledError not retried (1 attempt) Defaults favour retrying; configure them toward correctness.

2. Re-raise the original exception

By default, exhausted retries raise tenacity.RetryError, which wraps the last attempt. Callers that catch httpx.HTTPError or a domain exception miss it:

try:
    await get_profile(client, 42)
except httpx.HTTPError:
    ...                                  # without reraise=True this never runs: RetryError is raised


@retry(..., reraise=True)                # tested: the original ConnectionError propagates
async def get_profile(client: httpx.AsyncClient, user_id: int) -> dict: ...

Tested: without reraise=True, the caller received RetryError with the ConnectionError inside; with it, the ConnectionError itself. reraise=True keeps retries an implementation detail of the function — callers handle the same exceptions whether or not it retried. If you need the attempt count in the error, add it as a note in a before_sleep or retry_error_callback hook rather than changing the exception type.

Verify: exceptions reaching callers after exhausted retries are the same types the unwrapped call raises.

3. Bound total time with asyncio.timeout, not stop_after_delay

stop_after_delay is checked after each failed attempt, so the last attempt can run past the limit by its full duration. Measured: stop_after_delay(1.0) with attempts that timed out at 0.3 s ran for 1.30 s. An asyncio.timeout around the whole retry loop is a real deadline:

from tenacity import AsyncRetrying, retry_if_exception_type, wait_exponential_jitter


async def fetch_with_budget(client, url: str, budget: float = 1.0, per_attempt: float = 0.3):
    async with asyncio.timeout(budget):                          # hard deadline for everything
        async for attempt in AsyncRetrying(
            retry=retry_if_exception_type((TimeoutError, httpx.TransportError)),
            wait=wait_exponential_jitter(initial=0.05, max=0.2),
            reraise=True,
        ):
            with attempt:
                async with asyncio.timeout(per_attempt):         # bound each attempt
                    return await client.get(url)

Measured: the same three attempts ended at exactly 1.00 s. tenacity 8.3+ also offers stop_before_delay, which refuses to start an attempt that could not finish in time, but it cannot cut short an attempt in progress; the outer timeout can. The budget arithmetic for per-attempt and total timeouts is in setting per-attempt and total timeouts for retries.

Verify: with a dependency that never answers, the call fails at the budget, not at the budget plus one attempt.

Total time for a 1 s retry budget, attempts timing out at 0.3 s 2 horizontal bars comparing stop_after_delay(1.0) with the others. Total time for a 1 s retry budget, attempts timing out at 0.3 s stop_after_delay(1.0) 1.30 s, 3 attempts asyncio.timeout(1.0) around the loop 1.00 s, 3 attempts tenacity 9.1.4, Python 3.14; backoff 0.05-0.2 s. stop conditions run between attempts; a timeout runs during them.

4. Let cancellation through

Retry logic must never turn a cancellation into another attempt. tenacity handles this correctly for coroutines:

task = asyncio.create_task(get_profile(client, 42))
await asyncio.sleep(0.2)
task.cancel()            # tested: CancelledError immediately, even mid-backoff

Tested: cancelling during a long backoff sleep raised CancelledError at once, after one attempt; a wrapped call that itself raised CancelledError was not retried, because tenacity's retry conditions match Exception, and CancelledError is a BaseException. Keep it that way in your own predicates — never write retry_if_exception_type(BaseException) — and do not catch CancelledError inside the retried function. The same applies to the asyncio.timeout cancellation inside an attempt, which arrives as TimeoutError and is retryable by choice.

Verify: a test cancels a call during its backoff and asserts it finishes promptly with CancelledError.

5. Observe retries

Silent retries hide degraded dependencies until they fail outright. Log and count every retry:

import logging

from tenacity import before_sleep_log

log = logging.getLogger(__name__)


def count_retry(retry_state) -> None:
    RETRIES.labels(fn=retry_state.fn.__name__,
                   error=type(retry_state.outcome.exception()).__name__).inc()


@retry(
    retry=retry_if_exception(is_transient),
    stop=stop_after_attempt(4),
    wait=wait_exponential_jitter(initial=0.1, max=2.0),
    before_sleep=lambda rs: (count_retry(rs), before_sleep_log(log, logging.WARNING)(rs)),
    reraise=True,
)
async def get_profile(client: httpx.AsyncClient, user_id: int) -> dict: ...

before_sleep runs before each backoff with the outcome of the failed attempt. A retry counter per function and error type shows a dependency degrading long before error rates rise; tested, the decorated function's statistics also records attempt numbers and total idle time for debugging. Combine retries with a breaker and a budget so a real outage does not multiply load, as in combining circuit breakers with retries.

Verify: a dashboard shows retries per function and error type, and they correlate with dependency incidents.

How should this tenacity decorator be configured? A decision on What must the retry guarantee with 4 outcomes. How should this tenacity decorator be configured? What must the retry guarantee? only transient errors retry=retry_if_exception(pred) default retries all callers see real exceptions reraise=True not RetryError a hard deadline asyncio.timeout around it 1.00 vs 1.30 s visibility before_sleep: log + count early warning Four settings separate a safe retry from a load multiplier.

Verification

tenacity is configured safely when:

  • Every decorator has a retry= predicate limited to transient errors.
  • reraise=True keeps exception types unchanged for callers.
  • Total time is bounded by asyncio.timeout, with per-attempt timeouts inside.
  • Retries are logged and counted, and cancellation is never retried.

Diagnostic Hook: export retries and final failures per function. A rising retry rate with stable final failures means retries are absorbing a degrading dependency — investigate it now; final failures that equal attempts × calls mean the predicate is retrying errors that never succeed.

Pitfalls & edge cases

  • No retry= condition. Tested: ValueError retried three times.
  • Callers catching the original type. Without reraise=True they get RetryError.
  • stop_after_delay as a deadline. Tested: 1.30 s for a 1.0 s limit.
  • Retrying BaseException. Cancellation would be retried.

Frequently Asked Questions

Does tenacity work with async functions?

Yes. @retry on an async def retries with asyncio.sleep between attempts, and AsyncRetrying supports async for attempt loops. Tested with tenacity 9.1.4 on Python 3.14.

Why does tenacity raise RetryError instead of my exception?

That is the default after the last attempt. Set reraise=True to re-raise the original exception from the final attempt.

Does tenacity retry asyncio.CancelledError?

No. Its retry conditions match Exception, and CancelledError is a BaseException; in testing, a cancelled call and a call raising CancelledError were not retried.

How do I set a total timeout for tenacity retries?

Wrap the retry loop in async with asyncio.timeout(budget). stop_after_delay only checks between attempts and let a 1.0 s limit run to 1.30 s in testing.