Health Checking grpc.aio Services¶
gRPC has a standard health-checking protocol, grpc.health.v1.Health, and every gRPC-aware load balancer, service mesh and Kubernetes probe speaks it. Implementing it in a grpc.aio server is a few lines with the grpcio-health-checking package, and the interesting part is deciding what status to report and when to change it. Tested with grpcio 1.84: Check for a registered service returned SERVING; for an unregistered name it failed with NOT_FOUND; a Watch stream delivered SERVING immediately and NOT_SERVING within milliseconds of the status being changed; and after enter_graceful_shutdown(), the overall status read NOT_SERVING. A Check round trip over a local channel took about 221 µs, cheap enough to probe every second. This guide wires the service in, drives it from real dependency state, and connects it to Kubernetes and shutdown.
Prerequisites¶
- Python 3.11+,
pip install grpcio grpcio-health-checking; measured with grpcio 1.84. - A grpc.aio server, from building async gRPC services with grpc.aio.
- Probe design, from implementing health and readiness probes for asyncio.
1. Register the health service¶
Add the async HealthServicer to the same server as your services and set an initial status for each service name plus the empty name, which means "the server as a whole":
import grpc
from grpc_health.v1 import health, health_pb2, health_pb2_grpc
SERVING = health_pb2.HealthCheckResponse.SERVING
NOT_SERVING = health_pb2.HealthCheckResponse.NOT_SERVING
async def serve() -> None:
server = grpc.aio.server()
greet_pb2_grpc.add_GreeterServicer_to_server(Greeter(), server)
health_servicer = health.aio.HealthServicer()
health_pb2_grpc.add_HealthServicer_to_server(health_servicer, server)
server.add_insecure_port("[::]:50051")
await server.start()
await health_servicer.set("demo.Greeter", NOT_SERVING) # not ready until warmed
await health_servicer.set("", NOT_SERVING)
await warm_up()
await health_servicer.set("demo.Greeter", SERVING)
await health_servicer.set("", SERVING)
await server.wait_for_termination()
Use the fully qualified service name from the .proto (package.Service) so clients and load balancers can ask about a specific service. Unknown names are answered with the NOT_FOUND status code — tested — so a client that asks about a misspelled service gets an error, not a false "healthy". Start in NOT_SERVING and switch after warm-up, the same readiness rule as in warming connection pools at startup.
Verify: grpc_health_probe -addr=localhost:50051 -service=demo.Greeter reports SERVING once warm-up finishes.
2. Check and watch from clients¶
Clients call Check for a one-off answer, or Watch to receive a stream of status changes:
async with grpc.aio.insecure_channel("localhost:50051") as channel:
stub = health_pb2_grpc.HealthStub(channel)
reply = await stub.Check(health_pb2.HealthCheckRequest(service="demo.Greeter"), timeout=1.0)
print(health_pb2.HealthCheckResponse.ServingStatus.Name(reply.status)) # SERVING
async for update in stub.Watch(health_pb2.HealthCheckRequest(service="demo.Greeter")):
on_status_change(update.status) # first message: current status, then changes
Measured: 1,000 sequential Check calls over a local channel averaged 221 µs each. A Watch stream sent SERVING immediately and NOT_SERVING 0.2 s later, the moment the server changed it — no polling interval involved. Always pass a timeout to Check; a server that is down fails fast with UNAVAILABLE, but one that is overloaded may not answer at all.
Verify: a client watching the service sees a status change as soon as the server sets it.
3. Drive status from real dependency state¶
A health service that always says SERVING is a liveness check pretending to be readiness. Update the status from what the service actually needs — but only for dependencies whose failure means the service cannot do useful work:
async def health_loop(health_servicer, db_pool, interval: float = 2.0) -> None:
current = None
while True:
try:
async with asyncio.timeout(1.0):
await db_pool.fetchval("select 1")
status = SERVING
except Exception:
status = NOT_SERVING
if status != current: # only push changes
await health_servicer.set("demo.Greeter", status)
current = status
await asyncio.sleep(interval)
Setting the status only on change keeps Watch streams quiet and makes every pushed update meaningful. Be careful which dependencies count: if every instance reports NOT_SERVING when a shared database is down, the load balancer removes all of them and clients get connection errors instead of fast, informative failures. Reserve NOT_SERVING for problems local to the instance — its own pool broken, warm-up incomplete, shutting down — and let shared outages surface as errors from the RPCs themselves. This trade-off is discussed in Circuit Breakers & Bulkheads.
Verify: breaking the instance's own database pool flips its status within one interval; a database outage affecting all instances does not remove every instance at once, if that is your policy.
4. Hook it into Kubernetes probes¶
Kubernetes has native gRPC probes (stable since 1.27) that call Check:
readinessProbe:
grpc:
port: 50051
service: demo.Greeter # omit for the overall ("") status
periodSeconds: 2
failureThreshold: 2
livenessProbe:
grpc:
port: 50051 # overall status; keep this one simple
periodSeconds: 10
failureThreshold: 3
Use the specific service name for readiness, which reflects warm-up and local dependencies, and the overall status for liveness, which should only fail when the process is genuinely stuck. A liveness probe that fails on dependency problems restarts healthy containers in a loop. At about 221 µs per check, a 2-second readiness period costs nothing measurable. Older clusters use the grpc_health_probe binary as an exec probe with the same semantics.
Verify: kubectl describe pod shows the readiness probe passing after warm-up and failing when the status is set to NOT_SERVING.
5. Report NOT_SERVING during shutdown¶
On shutdown, flip every status to NOT_SERVING first, give load balancers time to notice, then stop the server with a grace period for in-flight RPCs:
import signal
async def serve() -> None:
...
stop = asyncio.Event()
loop = asyncio.get_running_loop()
loop.add_signal_handler(signal.SIGTERM, stop.set)
await stop.wait()
await health_servicer.enter_graceful_shutdown() # every service -> NOT_SERVING, now and for later sets
await asyncio.sleep(5) # let probes and balancers see it
await server.stop(grace=20) # finish in-flight RPCs, then close
Tested: after enter_graceful_shutdown(), the overall status read NOT_SERVING, and it ignores later attempts to set SERVING — so the health loop from step 3 cannot accidentally put the instance back into rotation. The sleep must exceed the readiness probe period times its failure threshold, so the orchestrator has removed the instance before it stops accepting RPCs. The full sequence is in propagating gRPC deadlines and cancellation and Graceful Shutdown & Signals.
Verify: during a rolling deploy, clients see no UNAVAILABLE errors from the instances being replaced.
Verification¶
Health checking is set up well when:
- Each service name and the overall name have a status, starting at
NOT_SERVING. - Status changes follow real local state, pushed only on change.
- Kubernetes readiness uses the service status; liveness uses the overall one.
- Shutdown flips to
NOT_SERVINGbefore the server stops.
Diagnostic Hook: log every status transition with its reason, and count Watch subscribers. Frequent flapping between SERVING and NOT_SERVING means the health loop's checks are too sensitive — add a failure threshold. Transitions that happen on every instance at the same moment mean a shared dependency is driving status, which removes the whole fleet at once.
Pitfalls & edge cases¶
- Always SERVING. The health service then says nothing useful.
- Shared dependencies in readiness. One outage removes every instance.
- Dependency checks in liveness. Healthy containers restart in a loop.
- Stopping the server before reporting NOT_SERVING. Clients see
UNAVAILABLEduring deploys.
Frequently Asked Questions¶
How do I add health checks to a grpc.aio server?
Install grpcio-health-checking, add health.aio.HealthServicer() to the server with health_pb2_grpc.add_HealthServicer_to_server, and call await servicer.set(service_name, status) for each service and for the empty name.
What does the gRPC health service return for an unknown service?
Check fails with the NOT_FOUND status code, as tested with grpcio 1.84, rather than returning a serving status.
How do Kubernetes probes check gRPC health?
Kubernetes has native grpc probes that call the standard Health Check method on a port, optionally for a specific service name; older clusters use the grpc_health_probe binary.
How should a gRPC server report health during shutdown?
Call enter_graceful_shutdown() to set every service to NOT_SERVING, wait longer than the readiness probe needs to notice, then call server.stop with a grace period.
Related¶
- gRPC & RPC — up to the topic overview.
- Load balancing grpc.aio clients — clients that act on health status.
- Network I/O & Protocol Handling — the section overview.