Making Retry Loops Cancellation-Safe¶
A retry loop is a place where exceptions are caught on purpose, which makes it the easiest place to catch the one exception that must never be caught: CancelledError. A loop that swallows it turns a cancellation — a request timeout, a client disconnect, a shutdown — into another retry. Measured on Python 3.14 with a call that failed after 100 ms and a loop allowing ten attempts with 100 ms backoff, cancelling the task at 0.45 s: a loop catching only ConnectionError ended with CancelledError at 0.45 s, as did tenacity's AsyncRetrying. A loop with a bare except: caught the cancellation, kept retrying until 1.95 s, and returned normally — the caller never learned it had been cancelled. A loop catching BaseException and re-raising only on the last attempt ran until 1.85 s and raised ConnectionError instead of CancelledError. A loop that slept with time.sleep for its backoff blocked the event loop, so the cancellation could not even be issued until 0.70 s. And retrying on TimeoutError from a per-attempt timeout was safe: an outer 1-second deadline still ended the loop at 1.00 s after 3 attempts. This guide writes retry loops that respect cancellation.
Prerequisites¶
- Python 3.11+ asyncio; tenacity optional.
- Retry basics, from writing a retry decorator for coroutines.
- The topic overview, Retry & Backoff Strategies.
1. Catch only the errors you mean to retry¶
asyncio.CancelledError derives from BaseException, not Exception, precisely so that except Exception does not catch it. Retry on the specific errors that a retry can fix:
RETRYABLE = (ConnectionError, TimeoutError, httpx.TransportError)
async def call_with_retry(op, attempts=10, base=0.1):
for attempt in range(1, attempts + 1):
try:
return await op()
except RETRYABLE:
if attempt == attempts:
raise
await asyncio.sleep(base * 2 ** (attempt - 1))
Measured with the task cancelled at 0.45 s — during its third attempt — the loop ended with CancelledError at 0.45 s. The cancellation arrived at the await inside op() or inside asyncio.sleep, passed through the except that did not match it, and left the loop. Narrow lists also stop retries on programming errors: a TypeError will not be fixed by trying again.
Verify: a test that cancels the retrying task mid-attempt and mid-backoff gets CancelledError both times, immediately.
2. Never use bare except or BaseException¶
A bare except: and except BaseException: both catch CancelledError. Measured with a bare except:, the cancelled task carried on retrying for another 1.5 seconds and then returned "gave up" as an ordinary value; the caller's await task returned instead of raising. With except BaseException: and a re-raise on the last attempt, the task ran on to 1.85 s and raised the final ConnectionError — the cancellation had been converted into a different error:
# Both of these swallow cancellation
async def bare(op, backoff):
try:
return await op()
except: # catches CancelledError, KeyboardInterrupt, SystemExit
await asyncio.sleep(backoff)
async def broad(op):
try:
return await op()
except BaseException: # same problem, more explicit
...
The damage reaches beyond the loop: a TaskGroup waiting for this task to finish after cancelling it waits the extra 1.5 seconds; asyncio.timeout around it does not fire on time; and shutdown hangs while the loop exhausts its attempts. Linters flag bare except: (ruff's E722, BLE001 for broad catches); enable them for async code.
Verify: a search for except: and except BaseException in retry code finds nothing, and the linter rule is on.
3. Back off with asyncio.sleep, never time.sleep¶
The backoff is part of the loop, and a blocking sleep blocks everything:
async def attempt_with_backoff(op):
try:
return await op()
except RETRYABLE:
time.sleep(0.6) # wrong: the whole event loop stops for 0.6 s
await asyncio.sleep(0.6) # right: other tasks run, cancellation can arrive
Measured with time.sleep(0.6) as the backoff: the code that was meant to cancel the task at 0.45 s could not run until the blocking sleep ended, and issued the cancellation at 0.70 s. Every other task in the process was frozen for the same time. A cancellation delivered during asyncio.sleep interrupts the backoff at once, which is the behaviour step 1 relies on.
Verify: no retry helper calls time.sleep, and a loop-lag probe stays flat while retries back off.
4. Retry per-attempt timeouts, not the overall deadline¶
Retrying on TimeoutError is common — an attempt that timed out may succeed on the next try. It is safe with asyncio.timeout, because each timeout converts only its own cancellation:
async def fetch_with_attempt_timeouts(op):
for attempt in range(10):
try:
async with asyncio.timeout(0.4): # per attempt
return await op()
except TimeoutError: # this attempt's timeout only
await asyncio.sleep(0.05)
async with asyncio.timeout(1.0): # the request's overall deadline
await fetch_with_attempt_timeouts(slow_op) # slow_op takes 0.6 s
Measured with an operation taking 0.6 s: attempts timed out at 0.4 s and were retried; the outer deadline fired during the third attempt, arrived inside the loop as CancelledError — not matched by except TimeoutError — and became TimeoutError at the outer block, at 1.00 s. The loop retried its own timeouts and not the caller's. This holds for asyncio.timeout and asyncio.wait_for on Python 3.11 and later; the distinction is covered in telling TimeoutError apart from CancelledError.
Verify: a test with per-attempt timeouts under a shorter outer deadline ends exactly at the outer deadline.
5. Use a library that gets this right, and test it¶
tenacity's AsyncRetrying retries on exceptions derived from Exception by default, so cancellation passes through — measured, it ended at 0.45 s like the hand-written narrow loop. Restrict it further to retryable errors, and make sure nothing inside the attempt swallows cancellation:
retrying = tenacity.AsyncRetrying(
retry=tenacity.retry_if_exception_type(RETRYABLE),
stop=tenacity.stop_after_attempt(5) | tenacity.stop_after_delay(10),
wait=tenacity.wait_random_exponential(multiplier=0.1, max=2),
reraise=True,
)
async def fetch(url):
async for attempt in retrying:
with attempt:
return await client.get(url)
Whatever the implementation, test cancellation explicitly — mid-attempt and mid-backoff — because the bug is invisible in normal operation. For deterministic tests of retry timing without real sleeps, see testing retry logic without real sleeps.
Verify: the retry helper has tests that cancel it in each phase and assert CancelledError within milliseconds.
Verification¶
Retry loops are cancellation-safe when:
- Only named retryable exceptions are caught — never bare
except:orBaseException. - Backoff uses
asyncio.sleep, so cancellation can interrupt it. - Per-attempt timeouts sit inside the caller's deadline, which is never retried.
- Tests cancel the loop mid-attempt and mid-backoff and expect
CancelledErrorat once.
Diagnostic Hook: when cancelling a request or shutting down a service takes seconds longer than expected, look for retry loops with except: or except BaseException. A bare except: kept a cancelled task retrying for 1.5 s and then returned normally in this test.
Pitfalls & edge cases¶
- Bare
except:in a retry loop. Measured: cancellation lost, task returned at 1.95 s. except BaseExceptionwith re-raise on last attempt. Measured:ConnectionErrorinstead of cancellation.time.sleepfor backoff. Measured: the cancel itself delayed to 0.70 s.- Retrying on
Exceptionbroadly. Programming errors get retried too.
Frequently Asked Questions¶
Does catching Exception catch asyncio.CancelledError?
No, since Python 3.8 CancelledError derives from BaseException. A loop catching specific Exception types stopped at once when cancelled.
Why does my cancelled task keep running?
A bare except: or except BaseException in a retry loop catches the cancellation. One such loop kept retrying for 1.5 s and returned normally.
Is it safe to retry on TimeoutError inside asyncio.timeout?
Yes: per-attempt timeouts were retried while an outer 1 s deadline still ended the loop at 1.00 s, because the outer timeout arrives inside as CancelledError.
Does tenacity handle asyncio cancellation correctly?
By default, yes: AsyncRetrying retries Exception subclasses, and a cancelled task ended at 0.45 s with CancelledError.
Related¶
- Retry & Backoff Strategies — up to the topic overview.
- Retrying database serialization failures — a retry loop that must also respect deadlines.
- Resilience, Cancellation & Error Handling — the section overview.