Serving gRPC and HTTP from One asyncio Process¶
Many services expose both gRPC and HTTP — gRPC for other services, HTTP for health checks, metrics, webhooks or a browser-facing API. Running both on one event loop is a few lines of code and shares state for free, but it couples their shutdown and their latency. Measured on Python 3.14 with grpcio 1.84.0, uvicorn 0.54 and Starlette 1.7: starting both with asyncio.gather(server.serve(), grpc_server.wait_for_termination()) and sending SIGTERM while 10 gRPC calls of 2 seconds and 10 HTTP requests of 1 second were in flight, the HTTP requests all returned 200, but all 10 gRPC calls failed with UNAVAILABLE and the process died by the signal, exit code -15. Handling the signal once and stopping both servers together completed all 20 and exited with code 0 in 0.27 s. Under load, sharing cost latency: with 32 clients hammering an HTTP endpoint at 7,039 requests per second, gRPC calls in the same process went from a 1.41 ms median to 16.33 ms; with gRPC and HTTP in separate processes, from 1.39 ms to 2.07 ms. This guide builds the combined process correctly and shows when to split it.
Prerequisites¶
- grpcio and uvicorn with an ASGI app.
- A gRPC service, from building async gRPC services with grpc.aio.
- The topic overview, gRPC & RPC.
1. See what the naive version does on SIGTERM¶
Both servers can run as coroutines on one loop. The obvious version:
async def main():
grpc_server = grpc.aio.server()
pbg.add_RecsServicer_to_server(Recs(), grpc_server)
grpc_server.add_insecure_port("0.0.0.0:50051")
await grpc_server.start()
http = uvicorn.Server(uvicorn.Config(app, host="0.0.0.0", port=8000))
await asyncio.gather(http.serve(), grpc_server.wait_for_termination())
It serves both protocols correctly — until shutdown. Measured with 10 gRPC calls taking 2 seconds and 10 HTTP requests taking 1 second, SIGTERM sent 0.3 s in: uvicorn caught the signal, drained its own requests — all 10 returned 200 — and then re-raised the signal it had captured, killing the process with exit code -15. The gRPC server was never asked to stop; its 10 calls failed with UNAVAILABLE when the process died. uvicorn's Server.serve() manages signals for itself, and only for itself.
Verify: a SIGTERM test with calls in flight on both protocols shows every call completing, not just the HTTP ones.
2. Own the signal and stop both servers¶
Register a loop signal handler for SIGTERM and SIGINT before uvicorn starts, and on the signal tell both servers to finish:
async def main():
grpc_server = grpc.aio.server()
pbg.add_RecsServicer_to_server(Recs(), grpc_server)
grpc_server.add_insecure_port("0.0.0.0:50051")
await grpc_server.start()
http = uvicorn.Server(uvicorn.Config(app, host="0.0.0.0", port=8000))
http_task = asyncio.create_task(http.serve())
stop = asyncio.Event()
loop = asyncio.get_running_loop()
for sig in (signal.SIGTERM, signal.SIGINT):
loop.add_signal_handler(sig, stop.set) # registered before serve() runs
await stop.wait()
http.should_exit = True # uvicorn: stop accepting, drain
await asyncio.gather(http_task, grpc_server.stop(grace=10))
Measured: all 10 HTTP requests returned 200, all 10 gRPC calls completed with OK, both servers stopped 0.20 s after the signal and the process exited with code 0 after 0.27 s. The order matters, and uvicorn's source shows why: Server.serve() replaces the process's signal handlers with its own while it runs, records any signal it receives, and on exit restores the previous handlers and re-raises the recorded signal. When the previous handler is the default, the re-raised SIGTERM kills the process — step 1. When it is the loop's handler, registered first, the re-raised signal reaches that handler and does nothing further. grpc_server.stop(grace) stops accepting new calls and waits up to grace seconds for in-flight ones before cancelling them; should_exit gives uvicorn's own graceful drain. Choose the grace period to fit inside the orchestrator's termination window, as in shutting down asyncio pods in Kubernetes.
Verify: the SIGTERM test from step 1 completes every in-flight call on both protocols and exits with code 0.
3. Share state, not the critical path¶
The main reason to combine the servers is shared in-process state: a gRPC service and its HTTP health or admin endpoints can read the same objects without a network hop:
STATE = {"calls": 0, "ready": False}
class Recs(pbg.RecsServicer):
async def Get(self, request, context):
STATE["calls"] += 1
return await build_batch(request)
async def health(request):
status = 200 if STATE["ready"] else 503
return JSONResponse({"grpc_calls": STATE["calls"]}, status_code=status)
Both servers run on the same loop, so plain Python objects are safe to share without locks as long as no await sits between a read and a dependent write. Pools and clients — database, Redis, HTTP — can be shared too, created once before both servers start and closed after both stop. That ordering is the same as in closing pools cleanly on shutdown.
Verify: shared resources are created before either server starts and closed after both have stopped.
4. Measure the latency coupling¶
Both servers' Python code runs on one thread, so CPU spent on one protocol delays the other. Measured with gRPC calls returning 10 records, sent at 200 per second, while 32 HTTP clients requested a 200-item JSON endpoint as fast as they could:
for i in range(1000):
await asyncio.sleep(max(0, t0 + i / 200 - time.perf_counter()))
t = time.perf_counter()
await stub.Get(pb.Req(n=10))
latencies.append(time.perf_counter() - t)
In one process, gRPC latency went from 1.41 ms median and 9.02 ms p99 without HTTP load to 16.33 ms median and 33.68 ms p99 with it, while HTTP served 7,039 requests per second. With gRPC and HTTP in two processes, gRPC went from 1.39 ms to 2.07 ms median and from 6.61 ms to 11.67 ms p99, and HTTP served 8,250 per second. In the shared process, every gRPC call queued behind HTTP request handling on the same loop.
Verify: a load test on the busier protocol shows the other protocol's latency staying within its objective.
5. Decide what to combine¶
The measurements suggest a rule. Combine when the HTTP side is light — health, readiness, metrics, an occasional admin call — and benefits from in-process state: its load is too small to delay gRPC, and one process is simpler to deploy. Split when both protocols carry real traffic: run them as separate processes, or separate deployments, each with its own loop and CPU. A middle ground is one image with a command-line switch:
def cli():
role = sys.argv[1] # "grpc", "http" or "both"
asyncio.run({"grpc": serve_grpc, "http": serve_http, "both": serve_both}[role]())
The same code then runs combined in development and split in production, with shared state replaced by a shared backend — a database or Redis — where the split processes need it. For CPU-heavy handlers on either side, the coupling grows; see comparing context switch costs of threads and tasks for what a busy thread does to a loop's latency.
Verify: the deployment documents which role each process runs, and combined processes carry only light secondary traffic.
Verification¶
A combined gRPC and HTTP process is correct when:
- One signal handler stops both servers, and in-flight calls on both complete.
- The process exits with code 0 after a SIGTERM, within the platform's grace period.
- Shared resources are created before both servers and closed after both.
- Load on one protocol does not push the other past its latency objective — measured, not assumed.
Diagnostic Hook: when gRPC clients see UNAVAILABLE during every deploy while HTTP clients do not, check whether uvicorn's Server.serve() is gathered with the gRPC server. uvicorn drained its own requests and then re-raised the signal, which ended the process before the gRPC server was told to stop.
Pitfalls & edge cases¶
- Gathering
serve()with the gRPC server. Measured: 10 of 10 gRPC calls failed on SIGTERM. - Default signal handling around
serve(). uvicorn re-raises the signal it captured after it exits. - Heavy HTTP traffic beside latency-sensitive RPCs. Measured: gRPC median rose from 1.41 to 16.33 ms.
- Closing shared pools when the first server stops. Close them after both.
Frequently Asked Questions¶
Can I run grpc.aio and uvicorn in the same event loop?
Yes. Start the gRPC server, run uvicorn.Server.serve() as a task, handle SIGTERM yourself, then set should_exit and await grpc_server.stop(grace) together.
Why do gRPC calls fail when my combined server shuts down?
uvicorn's serve() handles SIGTERM, drains HTTP, then re-raises the signal. Gathered with the gRPC server, all 10 in-flight gRPC calls failed with UNAVAILABLE.
Does HTTP traffic slow down gRPC in the same process?
It did here: gRPC median latency rose from 1.41 to 16.33 ms with HTTP at 7,039 req/s, against 1.39 to 2.07 ms when they ran as separate processes.
When should gRPC and HTTP share a process?
When the HTTP side is light, such as health checks and metrics, and benefits from in-process state. Split them when both carry real traffic.
Related¶
- gRPC & RPC — up to the topic overview.
- Health-checking grpc.aio services — the gRPC-native alternative to an HTTP health endpoint.
- Network I/O & Protocol Handling — the section overview.