Skip to content

Combining Circuit Breakers with Retries

Retries and circuit breakers solve opposite problems: retries push through transient failures, breakers stop sending traffic when failure is not transient. Used together carelessly, they fight — retries multiply load on a dependency that is down, and a breaker in the wrong place lets them. Simulated over a 10-second outage at 100 requests per second, with up to three retries per request and backoff of 0.1, 0.2 and 0.4 s: without a breaker, the dead dependency received 3,890 attempts and each failed request took 700 ms to fail. With a breaker checked before every attempt, it received 6, and failures were immediate. With the breaker wrapped around the whole retry loop, it received 24 — each logical failure had already spent its four attempts before the breaker counted it. This guide puts the two in the right order and makes their limits agree.

Prerequisites

1. See how retries amplify an outage

Each request that fails is retried; during an outage every attempt fails, so every request becomes four:

async def call_with_retries(fn, attempts: int = 4, backoff=(0.1, 0.2, 0.4)):
    for attempt in range(attempts):
        try:
            return await fn()
        except TransientError:
            if attempt == attempts - 1:
                raise
            await asyncio.sleep(backoff[attempt])

# simulated 10 s outage at 100 req/s: 3,890 attempts reached the dead dependency,
# and every failed request waited 700 ms (0.1 + 0.2 + 0.4) before failing

Fourfold traffic arrives exactly when the dependency can least take it, which often turns a short outage into a long one: the dependency restarts into a flood of retries and falls over again. Callers pay too — 700 ms of backoff before learning what the first attempt already showed. Retry budgets limit the amplification, as in implementing retry budgets to prevent retry storms; a breaker stops it.

Verify: during an incident, compare attempts per second at the dependency with requests per second at your service; a ratio near the maximum attempts means retries are amplifying the outage.

Attempts reaching a dead dependency during a 10 s outage 3 horizontal bars comparing retries, no breaker with the others. Attempts reaching a dead dependency during a 10 s outage retries, no breaker 3,890 breaker around the retry loop 24 breaker checked per attempt 6 Virtual-time simulation; 3 retries with 0.1/0.2/0.4 s backoff; breaker opens after 5 failures, open 5 s. The breaker must see every attempt, not every logical call.

2. Check the breaker before every attempt

Put the breaker inside the retry loop, so each attempt asks permission and reports its result:

class BreakerOpen(Exception):
    """Fail-fast signal: not retryable."""


async def call(breaker, fn, attempts: int = 4):
    for attempt in range(attempts):
        if not breaker.allow():
            raise BreakerOpen()                       # stop retrying the moment it opens
        try:
            result = await fn()
        except TransientError:
            breaker.record_failure()
            if attempt == attempts - 1:
                raise
            await asyncio.sleep(backoff_with_jitter(attempt))
        else:
            breaker.record_success()
            return result

Simulated: 6 attempts reached the dependency during the outage — the five that opened the breaker plus one half-open probe — and failed requests returned immediately instead of after 700 ms. The breaker sees the true failure rate of attempts, which is what it should judge, and an open breaker short-circuits the remaining retries of every in-flight request. BreakerOpen must not itself be retried; classify it as a permanent failure for that request.

Verify: with the breaker open, a request makes zero attempts and fails in microseconds.

3. Understand the cost of the breaker outside the loop

Wrapping the breaker around the whole retry loop is a common shape, often because the retry decorator and the breaker decorator are stacked in that order:

@circuit_breaker                  # outer: counts one failure per *logical* call
@retry(attempts=4)                # inner: makes 4 attempts before the breaker hears about it
async def fetch_profile(user_id: int) -> dict:
    ...

# simulated: 24 attempts reached the dead dependency instead of 6,
# and each failure that opened the breaker had already taken 700 ms

The breaker opens after five logical failures, each of which made four attempts, so it takes 20 attempts and 3.5 seconds of backoff to learn what five attempts showed. Requests that started before it opened still run their full retry sequence. Decorator order matters: the breaker must be the inner one, or the retry logic must consult it per attempt as in step 2.

Verify: inspect the decorator order on every call path that has both; the breaker sits closest to the call.

The right nesting of retries and a breaker A flow of 5 stages. The right nesting of retries and a breaker request retry loop breaker.allow()? open -> fail fast attempt with timeout record result per attempt backoff + jitter then next attempt Outer loop retries; inner breaker judges each attempt.

4. Make timeouts, retries and the breaker agree

Three settings interact: the per-attempt timeout, the retry schedule, and the breaker's window. They should be chosen together:

# A request's total budget must cover the attempts it may make
PER_ATTEMPT_TIMEOUT = 0.5
BACKOFF = (0.1, 0.2, 0.4)
TOTAL_BUDGET = 2.0                       # >= 0.5 * 4 + 0.7 would be 2.7; budget wins

# The breaker's window should see several requests' worth of attempts
BREAKER_WINDOW = 2.0                     # seconds
BREAKER_MIN_REQUESTS = 20


async def call_with_budget(breaker, fn):
    async with asyncio.timeout(TOTAL_BUDGET):
        return await call(breaker, lambda: with_timeout(fn, PER_ATTEMPT_TIMEOUT))

A total budget smaller than the full retry schedule means later retries are cut off — which is correct: retrying after the caller has stopped waiting only adds load. The breaker's window should be long relative to a single request's retries, so it judges a population of attempts rather than one request's bad luck. The budget arithmetic is covered in setting per-attempt and total timeouts for retries.

Verify: the worst-case time of a request — every attempt timing out, full backoff — is within its total budget, or the budget truncates it deliberately.

5. Decide what a fail-fast request returns

Once the breaker fails requests immediately, the caller needs something better than an exception — and the retry layer must not treat that exception as transient:

RETRYABLE = (TransientError, TimeoutError, ConnectionError)        # BreakerOpen is not here


async def get_recommendations(user_id: int) -> list[dict]:
    try:
        return await call(recs_breaker, lambda: recs_client.fetch(user_id))
    except BreakerOpen:
        BREAKER_FALLBACKS.labels(dependency="recs").inc()
        return await popular_items_cached()                        # degraded but useful

Fallbacks are what make a breaker user-friendly: cached results, defaults, a reduced feature. Where no fallback exists, return a clear 503 with Retry-After set near the breaker's remaining open time, so clients do not retry into it. Count fallbacks per dependency; they measure how much degraded service users received, which is the real cost of an outage, as discussed in adding timeouts and fallbacks to degraded dependencies.

Verify: with the breaker forced open in a test, the endpoint returns its fallback quickly and no retries are attempted.

How should retries and a breaker be arranged? A decision on What is being decided with 4 outcomes. How should retries and a breaker be arranged? What is being decided? where the breaker sits before each attempt 6 vs 24 attempts is BreakerOpen retryable? never fail fast how long overall total budget over retries truncates late retries what the caller gets fallback or 503 + Retry-After degraded, not broken Retries handle blips; the breaker ends them when the blip is an outage.

Verification

Retries and the breaker cooperate when:

  • The breaker is consulted before every attempt and records every attempt's result.
  • BreakerOpen is classified as non-retryable.
  • A total budget bounds the retry sequence, consistent with the breaker's window.
  • Fail-fast requests return a fallback or a 503 with Retry-After.

Diagnostic Hook: export attempts per request (a histogram) and breaker state per dependency. During an outage, attempts per request should drop to zero once the breaker opens; if they stay at the retry maximum, the breaker is outside the retry loop or BreakerOpen is being retried.

Pitfalls & edge cases

  • Retries without a breaker. Simulated: 3,890 attempts into a 10 s outage.
  • Breaker outside the retry decorator. Simulated: 24 attempts and 700 ms per failure before opening.
  • Retrying BreakerOpen. The breaker no longer fails fast.
  • Budgets shorter than one attempt. Every request times out before its first try completes.

Frequently Asked Questions

Should the circuit breaker be inside or outside the retry loop?

Inside: check the breaker before each attempt and record each attempt's result. In simulation that let 6 attempts reach a dead dependency, against 24 with the breaker around the whole retry loop.

Do retries make outages worse?

Without limits, yes: with three retries, 3,890 attempts hit a dependency during a simulated 10 s outage at 100 requests per second. Breakers and retry budgets stop the amplification.

Should a request retry when the circuit breaker is open?

No. Treat BreakerOpen as a permanent failure for that request and return a fallback or a 503 with Retry-After.

How do retries, timeouts and circuit breakers fit together?

A total time budget bounds the whole call, a per-attempt timeout bounds each try, the breaker is checked before each attempt, and backoff with jitter separates attempts.