Load Balancing grpc.aio Clients¶
gRPC multiplexes every call over a long-lived HTTP/2 connection, which defeats connection-level load balancing: a client that opens one connection to a load balancer's address sends all of its calls to whichever backend that connection landed on. To spread calls, the client has to know the backends and choose between them per call. Tested with grpcio 1.84, a channel to three local servers using the default pick_first policy sent all 300 calls to the first server; with round_robin it sent 100, 100, 100. When one backend was stopped a third of the way through a run, round-robin moved its share to the other two with zero failed calls, and pick_first switched to a surviving server with zero failures too. This guide resolves multiple backends, selects the policy, and covers the Kubernetes setup that makes it work.
Prerequisites¶
- Python 3.11+,
pip install grpcio; measured with grpcio 1.84. - A grpc.aio client, from building async gRPC services with grpc.aio.
- Health checking, from health checking grpc.aio services.
1. See why connection-level balancing fails¶
An L4 load balancer (a Kubernetes ClusterIP Service, a TCP load balancer) balances connections. A gRPC channel opens one connection per backend address it knows about and keeps it open:
# One address -> one connection -> one backend, for the life of the channel
channel = grpc.aio.insecure_channel("greeter.default.svc.cluster.local:50051")
stub = greet_pb2_grpc.GreeterStub(channel)
# every call from this process goes to the same pod
Measured with the default pick_first policy and three backend addresses: all 300 calls went to the first address. With a single virtual IP in front, the effect is the same — whichever pod received the connection gets all of this client's traffic, and new pods added by autoscaling receive none from existing clients. Load ends up proportional to how many clients happened to connect to each pod, not to capacity.
Verify: count calls per backend pod under load; with one virtual IP, a few pods carry most of the traffic.
2. Give the channel every backend address¶
Client-side balancing needs the backend list. The target string's scheme decides how it is resolved:
# Static list, useful for tests and fixed fleets
channel = grpc.aio.insecure_channel("ipv4:10.0.0.11:50051,10.0.0.12:50051,10.0.0.13:50051")
# DNS returning several A records: a Kubernetes headless Service, Consul, etc.
channel = grpc.aio.insecure_channel(
"dns:///greeter-headless.default.svc.cluster.local:50051",
options=[("grpc.dns_min_time_between_resolutions_ms", 10_000)],
)
The dns:/// resolver returns every A record and re-resolves periodically and when connections fail, so new and removed backends are picked up. In Kubernetes that means a headless Service (clusterIP: None), whose DNS name returns the pod IPs directly rather than one virtual IP. For anything more dynamic — weights, zones, service meshes — xDS-based resolution (xds:///) delegates the list and policy to a control plane.
Verify: dig +short greeter-headless.default.svc.cluster.local returns one address per ready pod.
3. Select round_robin in the service config¶
With several addresses, choose the policy in the channel's service config:
import json
channel = grpc.aio.insecure_channel(
"dns:///greeter-headless.default.svc.cluster.local:50051",
options=[("grpc.service_config", json.dumps({"loadBalancingConfig": [{"round_robin": {}}]}))],
)
Measured: 300 calls split 100/100/100 across three servers. round_robin keeps a connection to every address and rotates calls among the ready ones. pick_first — the default — connects to the first address that works and uses only it, which is right when you want affinity, such as a client talking to a primary with standbys behind it. Other built-in policies exist (weighted_round_robin, which uses server-reported load, and policies delivered through xDS), but round_robin is the right default for stateless services.
Verify: per-pod request rates are within a few percent of each other under steady load.
4. Survive backend failures¶
The policy tracks connection state per backend. A backend that goes away is skipped:
async with grpc.aio.insecure_channel(target, options=round_robin) as channel:
stub = greet_pb2_grpc.GreeterStub(channel)
for i in range(300):
if i == 100:
await servers[50072].stop(None) # kill one backend mid-run
await stub.Hello(greet_pb2.HelloRequest(name="x"), timeout=1)
# measured: 50071 -> 135, 50072 -> 33, 50073 -> 132 calls; 0 errors
The stopped server's connection closed cleanly (it sent GOAWAY), so the policy marked it not ready before the next call was assigned to it. A backend that disappears without closing — a crashed node, a network partition — is detected more slowly, through keepalive pings or failed calls; enable client keepalive so dead connections are found in seconds:
options = [
("grpc.service_config", json.dumps({"loadBalancingConfig": [{"round_robin": {}}]})),
("grpc.keepalive_time_ms", 20_000), # ping idle connections every 20 s
("grpc.keepalive_timeout_ms", 5_000), # consider dead after 5 s without an ack
]
Combine this with a retry policy for UNAVAILABLE, as in retrying gRPC calls by status code, so a call that does land on a dying backend is retried — usually on another one.
Verify: kill a backend pod during a load test; error rates stay at zero or near it, and the remaining pods absorb the traffic.
5. Rebalance when backends are added¶
A client balances only across the addresses it knows. After a scale-up, existing clients learn about new pods only when they re-resolve, and with long-lived connections they may rarely need to. Two server-side settings keep traffic flowing to new backends:
server = grpc.aio.server(options=[
("grpc.max_connection_age_ms", 300_000), # ask clients to reconnect every 5 min
("grpc.max_connection_age_grace_ms", 30_000), # let in-flight calls finish
])
max_connection_age makes the server send GOAWAY after a connection's age passes the limit; the client reconnects and re-resolves, picking up new addresses. Five minutes is a common setting: frequent enough to rebalance after scaling, rare enough that reconnect cost is negligible. Alternatively, put a gRPC-aware L7 proxy (Envoy, Linkerd, Istio) in front, which balances per call on the clients' behalf and needs no client configuration.
Verify: after doubling the backend count, per-pod request rates even out within the connection-age window.
Verification¶
Client-side load balancing is working when:
- The channel resolves every backend, through a headless Service or another multi-address resolver.
round_robinis selected in the service config for stateless services.- Backend failures produce no client errors, with keepalive and an
UNAVAILABLEretry policy. - Servers cap connection age, so scale-ups rebalance.
Diagnostic Hook: chart requests per second per backend pod. A flat, even line is balancing; a few hot pods mean clients are using pick_first or a single virtual IP; new pods that stay cold after a scale-up mean clients are not re-resolving — check max_connection_age on the servers.
Pitfalls & edge cases¶
- Using a ClusterIP Service for gRPC. One connection, one pod per client.
- Forgetting the policy. Several addresses with
pick_firststill use one. - No keepalive. Silently dead backends are found only by failed calls.
- Never-ending connections. New backends receive no traffic from existing clients.
Frequently Asked Questions¶
Why does all my gRPC traffic go to one pod?
gRPC keeps one long-lived HTTP/2 connection per address, and a ClusterIP Service exposes one address, so each client sticks to one pod. Use a headless Service with a dns:/// target and the round_robin policy, or an L7 proxy.
How do I enable round robin load balancing in grpc.aio?
Pass a service config with {"loadBalancingConfig": [{"round_robin": {}}]} as the grpc.service_config channel option, and use a target that resolves to several addresses. In testing, 300 calls split 100/100/100.
What is the default gRPC load balancing policy?
pick_first, which connects to the first working address and sends every call there. In testing it sent all 300 calls to one of three servers.
How do gRPC clients pick up new backends after scaling?
By re-resolving DNS, which happens when connections close. Set grpc.max_connection_age_ms on servers so connections are recycled periodically.
Related¶
- gRPC & RPC — up to the topic overview.
- Testing grpc.aio services — running several in-process servers like the ones measured here.
- Network I/O & Protocol Handling — the section overview.