Retrying Failed Background Jobs with Backoff¶
A background job that fails should usually be tried again — the database was failing over, the API returned 503, the network blipped. Retrying immediately is the wrong default: the dependency that just failed gets hit again at once, by every failing job, which is how one slow minute turns into a retry storm. Job systems retry by deferring: the failed job goes back into the queue with a delay, freeing the worker for other work. In a test with arq 0.28 against Redis 7.4, a job that failed twice and succeeded on its third try — deferred 0.2 s, then 0.4 s — completed in a burst worker run of 0.66 s, its result recorded with job_try = 3. This guide builds that retry policy properly: backoff with jitter, classification of retryable errors, a cap, and a place for jobs that exhaust it.
Prerequisites¶
- Python 3.11+,
pip install arqand a Redis server; the concepts transfer to Celery (self.retry(countdown=...)) and taskiq. - Backoff maths, from exponential backoff with jitter in asyncio.
- Job systems compared, from choosing between Celery, arq and taskiq.
1. Retry by deferring, not by looping¶
Inside a job, a retry loop with asyncio.sleep holds a worker slot for the whole wait. Raising arq's Retry instead releases the slot and puts the job back with a delay:
# pip install arq
from arq import Retry
from arq.worker import func
async def sync_invoice(ctx, invoice_id: str) -> str:
try:
await billing_api.push(invoice_id)
except (TimeoutError, ConnectionError) as exc:
raise Retry(defer=backoff(ctx["job_try"])) from exc
return "synced"
class WorkerSettings:
functions = [func(sync_invoice, max_tries=6)]
redis_settings = RedisSettings(host="redis")
ctx["job_try"] is 1 on the first attempt and increases with each retry; max_tries caps the total. Verified: a job raising Retry on tries 1 and 2 succeeded on try 3, with the worker free to run other jobs during the deferrals.
The worker-slot cost of in-job sleeping is real: with max_jobs=10 and a dependency outage that makes every job sleep 30 s between attempts, ten stuck jobs block the whole worker, including jobs for healthy dependencies.
Verify: during a forced dependency failure, other job types keep completing on the same workers.
2. Back off exponentially with jitter¶
Fixed delays synchronise failures: a thousand jobs that failed together retry together. Exponential growth spreads attempts out over time, and random jitter spreads them within each step:
import random
def backoff(attempt: int, base: float = 1.0, cap: float = 300.0) -> float:
"""Full-jitter exponential backoff: uniform between 0 and min(cap, base * 2^attempt)."""
return random.uniform(0, min(cap, base * 2 ** attempt))
With base=1 and cap=300, attempt 1 waits up to 2 s, attempt 5 up to 32 s, attempt 9 and beyond up to 5 minutes. Full jitter — a uniform draw from zero to the ceiling — is the variant that spreads load best when many jobs fail at once. Six tries with this schedule cover about a minute of outage in the worst case, which matches most dependency incidents; longer outages are better handled by the final-failure path than by retrying for hours.
Verify: fail 1,000 jobs at once; retry timestamps are spread across each backoff window rather than clustered at its end.
3. Retry only what can succeed later¶
Not every exception deserves a retry. A 404 will be a 404 on every attempt; a validation error will not fix itself. Classify:
import httpx
RETRYABLE_STATUS = {408, 425, 429, 500, 502, 503, 504}
def is_retryable(exc: BaseException) -> bool:
if isinstance(exc, (TimeoutError, ConnectionError, httpx.TransportError)):
return True
if isinstance(exc, httpx.HTTPStatusError):
return exc.response.status_code in RETRYABLE_STATUS
return False
async def sync_invoice(ctx, invoice_id: str) -> str:
try:
await billing_api.push(invoice_id)
except Exception as exc:
if is_retryable(exc) and ctx["job_try"] < MAX_TRIES:
delay = retry_after(exc) or backoff(ctx["job_try"])
raise Retry(defer=delay) from exc
raise # permanent, or out of tries: fail now
return "synced"
retry_after() reads a Retry-After header when the dependency sends one — a 429 telling you exactly when to come back is better information than any backoff formula, as covered in handling 429 Retry-After responses in async clients. The broader classification is in classifying retryable errors in async clients.
Verify: a job that hits a 404 fails on its first try; one that hits a 503 is retried.
4. Make retries safe to repeat¶
A retry re-runs the whole job. If the first attempt did half its work — wrote a row, then failed calling the API — the second attempt does it again. Every retried job must be idempotent:
async def charge_order(ctx, order_id: str) -> str:
order = await db.get_order(order_id)
if order.charged_at is not None:
return "already charged" # an earlier try got this far
charge = await payments.charge(order.total, idempotency_key=f"order:{order_id}")
await db.mark_charged(order_id, charge.id)
return charge.id
The idempotency key passed to the payment API makes the external side safe; the charged_at check makes the local side cheap. The full set of techniques is in making background jobs idempotent.
Verify: kill a worker in the middle of a job; after the retry, the side effect happened exactly once.
5. Give exhausted jobs somewhere to go¶
After max_tries, the job fails for good. Do not let that be silent. Record it somewhere a person or a process will look:
async def sync_invoice(ctx, invoice_id: str) -> str:
try:
...
except Exception as exc:
if is_retryable(exc) and ctx["job_try"] < MAX_TRIES:
raise Retry(defer=backoff(ctx["job_try"])) from exc
await ctx["redis"].lpush("dead:sync_invoice",
json.dumps({"invoice": invoice_id, "error": repr(exc)}))
log.error("sync_invoice %s failed permanently after %d tries", invoice_id, ctx["job_try"])
raise
A dead-letter list with the input and the last error lets you replay jobs once the cause is fixed, and its length is an excellent alert signal. The in-process version of the same idea is in implementing a dead letter queue with asyncio.
Verify: exhaust a job's tries; its input and error appear in the dead-letter list, and an alert fires.
Verification¶
Job retries are sound when:
- Retries defer through the job system instead of sleeping in the worker.
- Delays grow exponentially with jitter, or follow
Retry-Afterwhen given. - Permanent errors fail at once; only transient ones are retried.
- Jobs are idempotent, and exhausted jobs land in a dead-letter list with their inputs.
Diagnostic Hook: export retries per job type and per error class, and the dead-letter list length. A spike of retries for one job type with one error class is a dependency incident; dead letters growing for a job type with varied errors usually means a bug in the job itself. Alert on dead-letter growth, not on individual retries.
Pitfalls & edge cases¶
- Retrying with
asyncio.sleepinside the job. Holds a worker slot for every wait. - No jitter. Synchronised retries recreate the load spike that caused the failure.
- Retrying non-idempotent work. Duplicated charges, emails and writes.
- Unbounded retries. A job retried forever never alerts anyone.
Frequently Asked Questions¶
How do I retry a failed arq job with a delay?
Raise arq.Retry(defer=seconds) from the job. The job is put back in the queue and runs again after the delay, with ctx["job_try"] incremented. Set max_tries on the function to cap attempts.
What backoff should background job retries use?
Exponential backoff with full jitter — a random delay between zero and min(cap, base × 2^attempt) — so that many jobs failing together do not retry together. Use a Retry-After header instead when the dependency sends one.
Which errors should a background job retry?
Transient ones: timeouts, connection errors, 5xx responses and 429 rate limits. Not validation errors, 404s or other failures that will fail identically on every attempt.
What should happen when a job runs out of retries?
Record it in a dead-letter list with its input and last error, log it, and alert on the list's growth, so the job can be investigated and replayed instead of disappearing.
Related¶
- Background Jobs & Task Queues — up to the topic overview.
- Implementing retry budgets to prevent retry storms — capping retries across all jobs at once.
- Concurrent Execution & Worker Patterns — the section overview.