Skip to content

Retrying gRPC Calls by Status Code

gRPC clients can retry failed calls without any retry code in your application: a retryPolicy in the channel's service config tells the gRPC core which status codes to retry, how many attempts to make, and how to back off. Tested with grpcio 1.84 against a server that failed twice with UNAVAILABLE and then succeeded: with no policy the client saw UNAVAILABLE immediately; with a policy of up to 4 attempts retrying UNAVAILABLE, the call succeeded on the third attempt in 0.32 s, with backoff gaps of 0.10 and 0.22 s. A call failing with INVALID_ARGUMENT was not retried. A server that always failed got exactly 4 attempts in 0.67 s. And with a 0.3 s deadline, the same failing call made only 2 attempts and returned DEADLINE_EXCEEDED, not UNAVAILABLE — the deadline bounds the whole retry sequence. This guide configures the policy, chooses which codes to retry, and adds application-level retries where the built-in mechanism stops.

Prerequisites

1. Configure a retryPolicy in the service config

The service config is a JSON document passed as a channel option (or delivered by the name resolver). A methodConfig entry applies to a whole service or individual methods:

import json

import grpc

service_config = json.dumps({
    "methodConfig": [{
        "name": [{"service": "demo.Greeter"}],            # all methods of the service
        "retryPolicy": {
            "maxAttempts": 4,                              # including the first; capped at 5
            "initialBackoff": "0.1s",
            "maxBackoff": "1s",
            "backoffMultiplier": 2,
            "retryableStatusCodes": ["UNAVAILABLE"],
        },
    }]
})

channel = grpc.aio.insecure_channel(
    "greeter.internal:50051",
    options=[("grpc.service_config", service_config)],
)

Measured against a server failing the first two attempts: success on attempt 3, 0.32 s total; the observed gaps (0.10 s, then 0.22 s) are the jittered exponential backoff. The retries happen inside the gRPC core, below Python, so application code — and any client interceptors — see a single call. Retries are enabled by default in grpcio; the option ("grpc.enable_retries", 0) turns them off, which the measurement used for the no-policy baseline.

Verify: count attempts on the server side; a flaky dependency shows several attempts per client call, and the client sees success.

Attempts per call with a 4-attempt retryPolicy 4 horizontal bars comparing flaky (2 x UNAVAILABLE) with the others. Attempts per call with a 4-attempt retryPolicy flaky (2 x UNAVAILABLE) 3, then OK in 0.32 s INVALID_ARGUMENT 1, not retryable always UNAVAILABLE 4, gave up in 0.67 s always UNAVAILABLE, 0.3 s deadline 2, DEADLINE_EXCEEDED grpcio 1.84; policy maxAttempts 4, initialBackoff 0.1 s, multiplier 2, retrying UNAVAILABLE only. The status code decides whether to retry; the deadline decides how long.

2. Retry only codes that mean "try again"

Which status codes are safe to retry depends on whether the server could have acted on the request:

RETRY_ALWAYS_SAFE = ["UNAVAILABLE"]              # not processed: connection or server not ready
RETRY_IF_IDEMPOTENT = ["DEADLINE_EXCEEDED",      # may have been processed
                       "ABORTED",                 # concurrency conflict; retry the whole operation
                       "RESOURCE_EXHAUSTED",      # rate limit; back off longer
                       "INTERNAL", "UNKNOWN"]     # server bug or crash mid-call
NEVER_RETRY = ["INVALID_ARGUMENT", "NOT_FOUND", "ALREADY_EXISTS", "PERMISSION_DENIED",
               "UNAUTHENTICATED", "FAILED_PRECONDITION", "OUT_OF_RANGE", "UNIMPLEMENTED"]

UNAVAILABLE is the canonical retryable code: the server or connection was not ready, so the request did not run. The second group may have run partially or fully, so retry them only for methods that are idempotent — reads, or writes with an idempotency key. The third group describe the request itself; retrying sends the same bad request again. Tested, INVALID_ARGUMENT produced exactly one attempt under a policy that listed only UNAVAILABLE. Use per-method methodConfig entries to give read methods a broader list than write methods.

Verify: for each method, the retryable codes in the config match whether the method is idempotent.

3. Let deadlines bound the retries

A gRPC deadline covers the entire call, including every retry attempt and backoff. That is the correct behaviour — the caller's time budget is not extended by retries — but it changes the error you see:

stub = greet_pb2_grpc.GreeterStub(channel)
try:
    reply = await stub.Hello(request, timeout=0.3)
except grpc.aio.AioRpcError as exc:
    # measured against an always-UNAVAILABLE server:
    #   no deadline pressure -> UNAVAILABLE after 4 attempts (0.67 s)
    #   timeout=0.3          -> DEADLINE_EXCEEDED after 2 attempts (0.30 s)
    handle(exc.code())

When the deadline expires during retries, the call fails with DEADLINE_EXCEEDED, hiding the UNAVAILABLE that caused it. Log the details string and attempt metadata where you can, and size deadlines so that at least two or three attempts fit: with a 0.1 s initial backoff, a 0.3 s deadline left room for only two. Deadline propagation from caller to callee is covered in propagating gRPC deadlines and cancellation.

Verify: for each method, the deadline is at least the expected latency times the attempts you want, plus the backoffs.

Retries inside a deadline 2 lanes over time. Retries inside a deadline no deadline try 1 wait try 2 wait try 3 wait try 4 0.3 s deadline try 1 wait try 2 wait DEADLINE_EXCEEDED time, not to scale (0.67 s without a deadline) → The deadline is the budget for all attempts, not each one.

4. Retry in application code when the policy cannot

The built-in policy cannot retry everything: it does not retry after the server has sent response headers or messages (for streaming calls that started), it cannot refresh credentials or switch to a different endpoint, and it cannot tell idempotent requests from non-idempotent ones within one method. For those cases, an interceptor or wrapper retries at the application level:

import random

RETRYABLE = {grpc.StatusCode.UNAVAILABLE, grpc.StatusCode.RESOURCE_EXHAUSTED}


async def call_with_retry(method, request, *, attempts: int = 3, deadline: float = 2.0):
    loop = asyncio.get_running_loop()
    end = loop.time() + deadline
    for attempt in range(attempts):
        remaining = end - loop.time()
        if remaining <= 0:
            raise TimeoutError("retry budget exhausted")
        try:
            return await method(request, timeout=remaining)      # each attempt gets what is left
        except grpc.aio.AioRpcError as exc:
            if exc.code() not in RETRYABLE or attempt == attempts - 1:
                raise
            await asyncio.sleep(min(0.1 * 2 ** attempt, 1.0) * random.uniform(0.5, 1.0))

Passing the remaining budget as each attempt's timeout keeps the overall deadline intact, the same rule as in setting per-attempt and total timeouts for retries. Do not stack this on top of a channel retryPolicy for the same codes — the attempts multiply (3 application attempts × 4 core attempts = 12 requests).

Verify: a test that injects RESOURCE_EXHAUSTED shows the expected number of attempts and respects the total budget.

5. Protect the server with retry throttling

Retries add load exactly when the server is struggling. gRPC's service config has a token-bucket throttle that stops retrying when too many calls are failing:

service_config = json.dumps({
    "methodConfig": [{
        "name": [{"service": "demo.Greeter"}],
        "retryPolicy": {
            "maxAttempts": 4, "initialBackoff": "0.1s", "maxBackoff": "1s",
            "backoffMultiplier": 2, "retryableStatusCodes": ["UNAVAILABLE"],
        },
    }],
    "retryThrottling": {
        "maxTokens": 10,          # bucket size per server name
        "tokenRatio": 0.1,        # each success refills 0.1 tokens; each failure costs 1
    },
})

Each failed call costs one token and each success returns tokenRatio; while the bucket is below half full, retries stop and failures surface directly. A short outage is still smoothed over by retries, but a sustained one does not get multiplied traffic. This is a client-side version of the circuit breaker in Circuit Breakers & Bulkheads.

Verify: during a sustained outage in a test, the server's request rate stays close to the client's call rate rather than several times it.

Should this gRPC failure be retried? A decision on What status code came back with 3 outcomes. Should this gRPC failure be retried? What status code came back? UNAVAILABLE retryPolicy request did not run DEADLINE_EXCEEDED, ABORTED, RESOURCE_EXHAUSTED only if idempotent longer backoff INVALID_ARGUMENT, NOT_FOUND, PERMISSION_DENIED never same answer again Retry what the server did not do, never what it rejected.

Verification

gRPC retries are configured well when:

  • A retryPolicy retries UNAVAILABLE for every service, with jittered backoff.
  • Broader codes are retried only for idempotent methods.
  • Deadlines leave room for the intended attempts, and errors after retries are logged with their cause.
  • Retry throttling or an application budget prevents retry storms.

Diagnostic Hook: compare server-side request counts with client-side call counts per method. A ratio near 1 in normal operation that climbs during incidents is retries working; a ratio that climbs to maxAttempts and stays there means a sustained outage is being multiplied — add throttling. Many DEADLINE_EXCEEDED errors on methods that should retry mean the deadline leaves no room for attempts.

Pitfalls & edge cases

  • Retrying non-retryable codes. INVALID_ARGUMENT will fail the same way every time.
  • Deadlines too short for retries. Measured: only 2 of 4 attempts fit in 0.3 s.
  • Stacking application and channel retries. Attempts multiply.
  • No throttling. A sustained outage receives several times the normal traffic.

Frequently Asked Questions

How do I enable automatic retries in grpc.aio?

Pass a service config JSON with a retryPolicy (maxAttempts, backoff settings, retryableStatusCodes) as the grpc.service_config channel option. In testing, a call that failed twice with UNAVAILABLE succeeded on the third attempt in 0.32 s.

Which gRPC status codes should be retried?

UNAVAILABLE always; DEADLINE_EXCEEDED, ABORTED, RESOURCE_EXHAUSTED, INTERNAL and UNKNOWN only for idempotent methods; never INVALID_ARGUMENT, NOT_FOUND, PERMISSION_DENIED or other request errors.

Do gRPC retries respect the call deadline?

Yes. The deadline covers all attempts and backoffs. With a 0.3 s deadline, a failing call made two attempts and returned DEADLINE_EXCEEDED instead of UNAVAILABLE.

How do I stop gRPC retries from overloading a failing server?

Add retryThrottling to the service config, a token bucket that disables retries while too many calls are failing, or use a circuit breaker in application code.