Skip to content

Retrying Failed Background Jobs with Backoff

A background job that fails should usually be tried again — the database was failing over, the API returned 503, the network blipped. Retrying immediately is the wrong default: the dependency that just failed gets hit again at once, by every failing job, which is how one slow minute turns into a retry storm. Job systems retry by deferring: the failed job goes back into the queue with a delay, freeing the worker for other work. In a test with arq 0.28 against Redis 7.4, a job that failed twice and succeeded on its third try — deferred 0.2 s, then 0.4 s — completed in a burst worker run of 0.66 s, its result recorded with job_try = 3. This guide builds that retry policy properly: backoff with jitter, classification of retryable errors, a cap, and a place for jobs that exhaust it.

Prerequisites

1. Retry by deferring, not by looping

Inside a job, a retry loop with asyncio.sleep holds a worker slot for the whole wait. Raising arq's Retry instead releases the slot and puts the job back with a delay:

# pip install arq
from arq import Retry
from arq.worker import func


async def sync_invoice(ctx, invoice_id: str) -> str:
    try:
        await billing_api.push(invoice_id)
    except (TimeoutError, ConnectionError) as exc:
        raise Retry(defer=backoff(ctx["job_try"])) from exc
    return "synced"


class WorkerSettings:
    functions = [func(sync_invoice, max_tries=6)]
    redis_settings = RedisSettings(host="redis")

ctx["job_try"] is 1 on the first attempt and increases with each retry; max_tries caps the total. Verified: a job raising Retry on tries 1 and 2 succeeded on try 3, with the worker free to run other jobs during the deferrals.

The worker-slot cost of in-job sleeping is real: with max_jobs=10 and a dependency outage that makes every job sleep 30 s between attempts, ten stuck jobs block the whole worker, including jobs for healthy dependencies.

Verify: during a forced dependency failure, other job types keep completing on the same workers.

A job that succeeds on its third try 3 lanes over time. A job that succeeds on its third try job attempts try 1 fails try 2 fails try 3 ok deferred in Redis wait 0.2 s wait 0.4 s worker slot runs other jobs runs other jobs time → Measured with arq 0.28: the burst run finished in 0.66 s with job_try = 3 on the result.

2. Back off exponentially with jitter

Fixed delays synchronise failures: a thousand jobs that failed together retry together. Exponential growth spreads attempts out over time, and random jitter spreads them within each step:

import random


def backoff(attempt: int, base: float = 1.0, cap: float = 300.0) -> float:
    """Full-jitter exponential backoff: uniform between 0 and min(cap, base * 2^attempt)."""
    return random.uniform(0, min(cap, base * 2 ** attempt))

With base=1 and cap=300, attempt 1 waits up to 2 s, attempt 5 up to 32 s, attempt 9 and beyond up to 5 minutes. Full jitter — a uniform draw from zero to the ceiling — is the variant that spreads load best when many jobs fail at once. Six tries with this schedule cover about a minute of outage in the worst case, which matches most dependency incidents; longer outages are better handled by the final-failure path than by retrying for hours.

Verify: fail 1,000 jobs at once; retry timestamps are spread across each backoff window rather than clustered at its end.

Maximum wait before each retry, base 1 s, cap 300 s 5 horizontal bars comparing after try 5 with the others. Maximum wait before each retry, base 1 s, cap 300 s after try 5 up to 32 s after try 4 up to 16 s after try 3 up to 8 s after try 2 up to 4 s after try 1 up to 2 s Ceiling = min(300, 1 x 2^attempt); the actual delay is uniform between 0 and the ceiling. With max_tries=6 there are five retries, at most 62 s of waiting in total — enough for most dependency blips.

3. Retry only what can succeed later

Not every exception deserves a retry. A 404 will be a 404 on every attempt; a validation error will not fix itself. Classify:

import httpx

RETRYABLE_STATUS = {408, 425, 429, 500, 502, 503, 504}


def is_retryable(exc: BaseException) -> bool:
    if isinstance(exc, (TimeoutError, ConnectionError, httpx.TransportError)):
        return True
    if isinstance(exc, httpx.HTTPStatusError):
        return exc.response.status_code in RETRYABLE_STATUS
    return False


async def sync_invoice(ctx, invoice_id: str) -> str:
    try:
        await billing_api.push(invoice_id)
    except Exception as exc:
        if is_retryable(exc) and ctx["job_try"] < MAX_TRIES:
            delay = retry_after(exc) or backoff(ctx["job_try"])
            raise Retry(defer=delay) from exc
        raise                                     # permanent, or out of tries: fail now
    return "synced"

retry_after() reads a Retry-After header when the dependency sends one — a 429 telling you exactly when to come back is better information than any backoff formula, as covered in handling 429 Retry-After responses in async clients. The broader classification is in classifying retryable errors in async clients.

Verify: a job that hits a 404 fails on its first try; one that hits a 503 is retried.

Should this job failure be retried? A decision on What did the job fail with with 3 outcomes. Should this job failure be retried? What did the job fail with? timeout, 5xx, connection reset Retry(defer=backoff) likely transient 429 with Retry-After Retry(defer=header) the server said when 4xx, validation, not found fail now retrying cannot help Retrying a permanent error only delays the alert and wastes worker time.

4. Make retries safe to repeat

A retry re-runs the whole job. If the first attempt did half its work — wrote a row, then failed calling the API — the second attempt does it again. Every retried job must be idempotent:

async def charge_order(ctx, order_id: str) -> str:
    order = await db.get_order(order_id)
    if order.charged_at is not None:
        return "already charged"                       # an earlier try got this far
    charge = await payments.charge(order.total, idempotency_key=f"order:{order_id}")
    await db.mark_charged(order_id, charge.id)
    return charge.id

The idempotency key passed to the payment API makes the external side safe; the charged_at check makes the local side cheap. The full set of techniques is in making background jobs idempotent.

Verify: kill a worker in the middle of a job; after the retry, the side effect happened exactly once.

5. Give exhausted jobs somewhere to go

After max_tries, the job fails for good. Do not let that be silent. Record it somewhere a person or a process will look:

async def sync_invoice(ctx, invoice_id: str) -> str:
    try:
        ...
    except Exception as exc:
        if is_retryable(exc) and ctx["job_try"] < MAX_TRIES:
            raise Retry(defer=backoff(ctx["job_try"])) from exc
        await ctx["redis"].lpush("dead:sync_invoice",
                                 json.dumps({"invoice": invoice_id, "error": repr(exc)}))
        log.error("sync_invoice %s failed permanently after %d tries", invoice_id, ctx["job_try"])
        raise

A dead-letter list with the input and the last error lets you replay jobs once the cause is fixed, and its length is an excellent alert signal. The in-process version of the same idea is in implementing a dead letter queue with asyncio.

Verify: exhaust a job's tries; its input and error appear in the dead-letter list, and an alert fires.

Verification

Job retries are sound when:

  • Retries defer through the job system instead of sleeping in the worker.
  • Delays grow exponentially with jitter, or follow Retry-After when given.
  • Permanent errors fail at once; only transient ones are retried.
  • Jobs are idempotent, and exhausted jobs land in a dead-letter list with their inputs.

Diagnostic Hook: export retries per job type and per error class, and the dead-letter list length. A spike of retries for one job type with one error class is a dependency incident; dead letters growing for a job type with varied errors usually means a bug in the job itself. Alert on dead-letter growth, not on individual retries.

Pitfalls & edge cases

  • Retrying with asyncio.sleep inside the job. Holds a worker slot for every wait.
  • No jitter. Synchronised retries recreate the load spike that caused the failure.
  • Retrying non-idempotent work. Duplicated charges, emails and writes.
  • Unbounded retries. A job retried forever never alerts anyone.

Frequently Asked Questions

How do I retry a failed arq job with a delay?

Raise arq.Retry(defer=seconds) from the job. The job is put back in the queue and runs again after the delay, with ctx["job_try"] incremented. Set max_tries on the function to cap attempts.

What backoff should background job retries use?

Exponential backoff with full jitter — a random delay between zero and min(cap, base × 2^attempt) — so that many jobs failing together do not retry together. Use a Retry-After header instead when the dependency sends one.

Which errors should a background job retry?

Transient ones: timeouts, connection errors, 5xx responses and 429 rate limits. Not validation errors, 404s or other failures that will fail identically on every attempt.

What should happen when a job runs out of retries?

Record it in a dead-letter list with its input and last error, log it, and alert on the list's growth, so the job can be investigated and replayed instead of disappearing.