Compressing gRPC Messages in grpc.aio¶
gRPC supports per-message compression, negotiated between client and server, and grpc.aio exposes it as one argument on the server, the channel or a single call. Whether to turn it on depends on what the messages contain and what the network costs, and both are easy to measure. Measured on Python 3.14 with grpcio 1.84.0 over a local channel, with a byte-counting TCP proxy between client and server: a response of 20,000 user records, 1,115 KiB uncompressed, went on the wire as 303 KiB with gzip or deflate — 27% of the size. On the server, CPU per call rose from about 6 ms to 23–24 ms, and on a local link the call took 31.6 ms instead of 13.7 ms. A 1 MiB response of random bytes stayed at 1,024 KiB on the wire with compression enabled, yet the server still spent 23–24 ms of CPU per call trying. Small single-record calls showed no measurable difference, 386–506 µs either way. This guide turns those numbers into a policy.
Prerequisites¶
- grpcio and a service built with
grpc.aio. - Service basics, from building async gRPC services with grpc.aio.
- The topic overview, gRPC & RPC.
1. Measure bytes on the wire¶
Compression ratios depend entirely on the payload, so measure on real messages. A small asyncio proxy that forwards traffic and counts bytes from the server gives the wire size without packet capture:
async def pump(reader, writer, count: bool):
while data := await reader.read(65536):
if count:
total[0] += len(data)
writer.write(data)
await writer.drain()
Measured over 50 calls each: the records response — 20,000 messages with an id, name, email, score and tags, 1,141,266 bytes serialized — was 1,115 KiB per response uncompressed and 303 KiB with either gzip or deflate. A 1 MiB random bytes field was 1,024 KiB in every case. Protobuf's binary encoding already removes field names, so text-like repeated values — names, emails, tags — are what compression finds to remove; already-compressed data such as images, archives or encrypted blobs gains nothing.
Verify: wire sizes for the service's largest and most frequent responses are measured with and without compression.
2. Enable compression where it helps¶
Compression can be set as the server's default, the channel's default for requests, or per call:
server = grpc.aio.server(compression=grpc.Compression.Gzip) # responses
channel = grpc.aio.insecure_channel(
"svc:50051", compression=grpc.Compression.Gzip, # requests
)
reply = await stub.Get(req, compression=grpc.Compression.Gzip) # this call's request only
A handler can exempt one response from the server's default, and this is where measurement mattered. In grpcio 1.84.0's asyncio server, context.set_compression(...) had no measurable effect in either direction: with the server defaulting to gzip, setting NoCompression for the random payload left server CPU at about 23 ms per call; with the server defaulting to none, setting Gzip for the records left them at 1,115 KiB on the wire. context.disable_next_message_compression() did work:
class Recs(pbg.RecsServicer):
async def Get(self, request, context):
if request.random:
context.disable_next_message_compression() # already incompressible
return await build_batch(request)
Measured with the server defaulting to gzip: the random response dropped to about 1 ms of server CPU and 5.98 ms per call, while the records response stayed compressed at 303 KiB. The practical policy is therefore a compressing server default with explicit exemptions — or separate servers or ports for compressible and incompressible services — verified by measuring wire bytes, not by trusting the API call. gzip and deflate produced the same 303 KiB and similar CPU here; gzip is the more widely supported choice across gRPC implementations.
Verify: a byte count on the wire confirms that compressible responses are compressed and exempted ones are not.
3. Weigh CPU against bandwidth¶
On a local link, compression made the call slower — 31.6 ms against 13.7 ms — because the network cost almost nothing and compression cost 17 ms of server CPU. On a constrained link, the arithmetic changes. Saving 812 KiB per response is worth about 66 ms at 100 Mbit/s and about 6.6 ms at 1 Gbit/s — a calculation from the measured sizes, not a measurement:
def transfer_ms(kib: float, mbit_per_s: float) -> float:
return kib * 1024 * 8 / (mbit_per_s * 1_000_000) * 1000
saved = transfer_ms(1115, 100) - transfer_ms(303, 100) # ≈ 66.5 ms at 100 Mbit/s
cost = 23 - 6 # extra server CPU, ms
Compression pays when the bandwidth saved exceeds the CPU spent: across regions, to mobile clients, over metered links, or when egress is billed by the byte. Within a data centre on 10 Gbit/s links, the same response spends more CPU than it saves time. Server CPU is also a capacity cost: at 23 ms per call, one core handles about 43 such responses per second, against about 160 uncompressed.
Verify: the decision is made per deployment path, using measured sizes and CPU and the real link speed.
4. Leave small messages alone¶
Most RPCs carry small messages, where the question is overhead rather than savings. Measured with a single-record response over 3,000 calls: 433 µs per call with no compression, 412 µs with the server defaulting to gzip, and 386–506 µs with compression switched per RPC — differences within run-to-run noise. Compressing a few hundred bytes neither helps nor hurts measurably, so the policy can focus on large responses. gRPC compresses each message separately, so in a server-streaming RPC each chunk is compressed on its own:
async def DownloadStream(self, request, context): # server default: gzip
async for chunk in read_chunks(request.path, size=64 * 1024):
if chunk_is_compressed(chunk):
context.disable_next_message_compression() # per message, as measured above
yield pb.Chunk(data=chunk)
Chunks of tens of KiB still compress well when their content is compressible; very small chunks compress poorly because each starts with an empty dictionary. For large transfers, combine streaming with compression, as in sending large messages with gRPC AsyncIO.
Verify: small RPCs are left at the server's default, and streamed chunks are large enough to compress.
5. Watch both sides after enabling it¶
Compression moves work: the sender spends CPU compressing, the receiver spends CPU decompressing. Measured on the client: 5.94 ms of CPU per records call uncompressed and 7.97–8.29 ms compressed — a smaller increase than the server's, since decompression is cheaper than compression. Record both:
start_cpu = time.process_time()
reply = await stub.Get(req)
rpc_cpu_ms.labels(method="Get").observe((time.process_time() - start_cpu) * 1000)
time.process_time() measures the whole process, so this is only an approximation in a busy service; per-RPC CPU is better taken from a load test that runs one RPC type at a time. After enabling compression in production, compare server CPU, p99 latency and egress bytes with the previous week; all three should move in the expected direction.
Verify: server CPU, latency and egress are compared before and after the change, and the change is reverted where CPU rose without a matching latency or cost benefit.
Verification¶
gRPC compression is configured well when:
- Wire sizes are measured for large responses, with and without compression.
- Incompressible responses are exempted with
disable_next_message_compression(), confirmed by wire bytes. - The CPU cost is justified by bandwidth or egress savings on the actual link.
- Small messages are left alone, since they showed no measurable difference.
Diagnostic Hook: when a gRPC server's CPU rises after enabling compression with no drop in egress, look at what the large responses contain. A 1 MiB random payload cost about 24 ms of compression CPU per call and left the wire size at 1,024 KiB.
Pitfalls & edge cases¶
- Compressing already-compressed data. Measured: no saving, about 24 ms of CPU per call.
- Judging compression on a local link. It made calls slower: 31.6 ms against 13.7 ms.
- Trusting
context.set_compression. It changed nothing measurable in grpcio 1.84.0's aio server. - Tiny streamed chunks. Each chunk is compressed on its own.
Frequently Asked Questions¶
How do I enable gzip compression in grpc.aio?
Pass compression=grpc.Compression.Gzip to grpc.aio.server for responses or to the channel for requests. To exempt one response, call context.disable_next_message_compression(); set_compression had no effect here.
How much does gRPC gzip compression save?
For a response of 20,000 user records, 1,115 KiB became 303 KiB. Random bytes stayed at 1,024 KiB. Measure your own payloads.
Does gRPC compression slow down calls?
On a fast link, yes: server CPU rose from about 6 to 23 ms per 1 MB response and the call took 31.6 ms instead of 13.7 ms locally. It pays on slow or metered links.
Is deflate better than gzip for gRPC?
They gave the same 303 KiB and similar CPU here. Gzip is the more widely supported choice across gRPC implementations.
Related¶
- gRPC & RPC — up to the topic overview.
- Serving gRPC and HTTP in one process — sharing one process between protocols.
- Network I/O & Protocol Handling — the section overview.