Skip to content

Load Balancing grpc.aio Clients

gRPC multiplexes every call over a long-lived HTTP/2 connection, which defeats connection-level load balancing: a client that opens one connection to a load balancer's address sends all of its calls to whichever backend that connection landed on. To spread calls, the client has to know the backends and choose between them per call. Tested with grpcio 1.84, a channel to three local servers using the default pick_first policy sent all 300 calls to the first server; with round_robin it sent 100, 100, 100. When one backend was stopped a third of the way through a run, round-robin moved its share to the other two with zero failed calls, and pick_first switched to a surviving server with zero failures too. This guide resolves multiple backends, selects the policy, and covers the Kubernetes setup that makes it work.

Prerequisites

1. See why connection-level balancing fails

An L4 load balancer (a Kubernetes ClusterIP Service, a TCP load balancer) balances connections. A gRPC channel opens one connection per backend address it knows about and keeps it open:

# One address -> one connection -> one backend, for the life of the channel
channel = grpc.aio.insecure_channel("greeter.default.svc.cluster.local:50051")
stub = greet_pb2_grpc.GreeterStub(channel)
# every call from this process goes to the same pod

Measured with the default pick_first policy and three backend addresses: all 300 calls went to the first address. With a single virtual IP in front, the effect is the same — whichever pod received the connection gets all of this client's traffic, and new pods added by autoscaling receive none from existing clients. Load ends up proportional to how many clients happened to connect to each pod, not to capacity.

Verify: count calls per backend pod under load; with one virtual IP, a few pods carry most of the traffic.

300 calls across three backends, by policy 3 horizontal bars comparing pick_first: backend 1 with the others. 300 calls across three backends, by policy pick_first: backend 1 300 pick_first: backends 2, 3 0 each round_robin: each backend 100 each grpcio 1.84; target ipv4:127.0.0.1:50071,50072,50073; sequential unary calls. The policy, not the number of addresses, decides how calls spread.

2. Give the channel every backend address

Client-side balancing needs the backend list. The target string's scheme decides how it is resolved:

# Static list, useful for tests and fixed fleets
channel = grpc.aio.insecure_channel("ipv4:10.0.0.11:50051,10.0.0.12:50051,10.0.0.13:50051")

# DNS returning several A records: a Kubernetes headless Service, Consul, etc.
channel = grpc.aio.insecure_channel(
    "dns:///greeter-headless.default.svc.cluster.local:50051",
    options=[("grpc.dns_min_time_between_resolutions_ms", 10_000)],
)

The dns:/// resolver returns every A record and re-resolves periodically and when connections fail, so new and removed backends are picked up. In Kubernetes that means a headless Service (clusterIP: None), whose DNS name returns the pod IPs directly rather than one virtual IP. For anything more dynamic — weights, zones, service meshes — xDS-based resolution (xds:///) delegates the list and policy to a control plane.

Verify: dig +short greeter-headless.default.svc.cluster.local returns one address per ready pod.

3. Select round_robin in the service config

With several addresses, choose the policy in the channel's service config:

import json

channel = grpc.aio.insecure_channel(
    "dns:///greeter-headless.default.svc.cluster.local:50051",
    options=[("grpc.service_config", json.dumps({"loadBalancingConfig": [{"round_robin": {}}]}))],
)

Measured: 300 calls split 100/100/100 across three servers. round_robin keeps a connection to every address and rotates calls among the ready ones. pick_first — the default — connects to the first address that works and uses only it, which is right when you want affinity, such as a client talking to a primary with standbys behind it. Other built-in policies exist (weighted_round_robin, which uses server-reported load, and policies delivered through xDS), but round_robin is the right default for stateless services.

Verify: per-pod request rates are within a few percent of each other under steady load.

How a balanced gRPC call finds a backend A flow of 4 stages. How a balanced gRPC call finds a backend resolver dns:/// -> pod IPs LB policy subchannel per address per call next ready subchannel backend fails skip, reconnect later Balancing happens per call, inside the client's channel.

4. Survive backend failures

The policy tracks connection state per backend. A backend that goes away is skipped:

async with grpc.aio.insecure_channel(target, options=round_robin) as channel:
    stub = greet_pb2_grpc.GreeterStub(channel)
    for i in range(300):
        if i == 100:
            await servers[50072].stop(None)                      # kill one backend mid-run
        await stub.Hello(greet_pb2.HelloRequest(name="x"), timeout=1)
# measured: 50071 -> 135, 50072 -> 33, 50073 -> 132 calls; 0 errors

The stopped server's connection closed cleanly (it sent GOAWAY), so the policy marked it not ready before the next call was assigned to it. A backend that disappears without closing — a crashed node, a network partition — is detected more slowly, through keepalive pings or failed calls; enable client keepalive so dead connections are found in seconds:

options = [
    ("grpc.service_config", json.dumps({"loadBalancingConfig": [{"round_robin": {}}]})),
    ("grpc.keepalive_time_ms", 20_000),         # ping idle connections every 20 s
    ("grpc.keepalive_timeout_ms", 5_000),       # consider dead after 5 s without an ack
]

Combine this with a retry policy for UNAVAILABLE, as in retrying gRPC calls by status code, so a call that does land on a dying backend is retried — usually on another one.

Verify: kill a backend pod during a load test; error rates stay at zero or near it, and the remaining pods absorb the traffic.

5. Rebalance when backends are added

A client balances only across the addresses it knows. After a scale-up, existing clients learn about new pods only when they re-resolve, and with long-lived connections they may rarely need to. Two server-side settings keep traffic flowing to new backends:

server = grpc.aio.server(options=[
    ("grpc.max_connection_age_ms", 300_000),        # ask clients to reconnect every 5 min
    ("grpc.max_connection_age_grace_ms", 30_000),   # let in-flight calls finish
])

max_connection_age makes the server send GOAWAY after a connection's age passes the limit; the client reconnects and re-resolves, picking up new addresses. Five minutes is a common setting: frequent enough to rebalance after scaling, rare enough that reconnect cost is negligible. Alternatively, put a gRPC-aware L7 proxy (Envoy, Linkerd, Istio) in front, which balances per call on the clients' behalf and needs no client configuration.

Verify: after doubling the backend count, per-pod request rates even out within the connection-age window.

Where should gRPC load balancing happen? A decision on What is in front of the backends with 3 outcomes. Where should gRPC load balancing happen? What is in front of the backends? nothing, clients you control dns:/// + round_robin headless Service Envoy, Istio, Linkerd proxy balances per call plain channel primary + standbys pick_first ordered addresses In all cases, servers should cap connection age so new backends get traffic.

Verification

Client-side load balancing is working when:

  • The channel resolves every backend, through a headless Service or another multi-address resolver.
  • round_robin is selected in the service config for stateless services.
  • Backend failures produce no client errors, with keepalive and an UNAVAILABLE retry policy.
  • Servers cap connection age, so scale-ups rebalance.

Diagnostic Hook: chart requests per second per backend pod. A flat, even line is balancing; a few hot pods mean clients are using pick_first or a single virtual IP; new pods that stay cold after a scale-up mean clients are not re-resolving — check max_connection_age on the servers.

Pitfalls & edge cases

  • Using a ClusterIP Service for gRPC. One connection, one pod per client.
  • Forgetting the policy. Several addresses with pick_first still use one.
  • No keepalive. Silently dead backends are found only by failed calls.
  • Never-ending connections. New backends receive no traffic from existing clients.

Frequently Asked Questions

Why does all my gRPC traffic go to one pod?

gRPC keeps one long-lived HTTP/2 connection per address, and a ClusterIP Service exposes one address, so each client sticks to one pod. Use a headless Service with a dns:/// target and the round_robin policy, or an L7 proxy.

How do I enable round robin load balancing in grpc.aio?

Pass a service config with {"loadBalancingConfig": [{"round_robin": {}}]} as the grpc.service_config channel option, and use a target that resolves to several addresses. In testing, 300 calls split 100/100/100.

What is the default gRPC load balancing policy?

pick_first, which connects to the first working address and sends every call there. In testing it sent all 300 calls to one of three servers.

How do gRPC clients pick up new backends after scaling?

By re-resolving DNS, which happens when connections close. Set grpc.max_connection_age_ms on servers so connections are recycled periodically.