Skip to content

Compressing gRPC Messages in grpc.aio

gRPC supports per-message compression, negotiated between client and server, and grpc.aio exposes it as one argument on the server, the channel or a single call. Whether to turn it on depends on what the messages contain and what the network costs, and both are easy to measure. Measured on Python 3.14 with grpcio 1.84.0 over a local channel, with a byte-counting TCP proxy between client and server: a response of 20,000 user records, 1,115 KiB uncompressed, went on the wire as 303 KiB with gzip or deflate — 27% of the size. On the server, CPU per call rose from about 6 ms to 23–24 ms, and on a local link the call took 31.6 ms instead of 13.7 ms. A 1 MiB response of random bytes stayed at 1,024 KiB on the wire with compression enabled, yet the server still spent 23–24 ms of CPU per call trying. Small single-record calls showed no measurable difference, 386–506 µs either way. This guide turns those numbers into a policy.

Prerequisites

1. Measure bytes on the wire

Compression ratios depend entirely on the payload, so measure on real messages. A small asyncio proxy that forwards traffic and counts bytes from the server gives the wire size without packet capture:

async def pump(reader, writer, count: bool):
    while data := await reader.read(65536):
        if count:
            total[0] += len(data)
        writer.write(data)
        await writer.drain()

Measured over 50 calls each: the records response — 20,000 messages with an id, name, email, score and tags, 1,141,266 bytes serialized — was 1,115 KiB per response uncompressed and 303 KiB with either gzip or deflate. A 1 MiB random bytes field was 1,024 KiB in every case. Protobuf's binary encoding already removes field names, so text-like repeated values — names, emails, tags — are what compression finds to remove; already-compressed data such as images, archives or encrypted blobs gains nothing.

Verify: wire sizes for the service's largest and most frequent responses are measured with and without compression.

50 calls per setting, local channel A grid of 6 rows by 6 columns. 50 calls per setting, local channel payload compression on the wire call time server CPU client CPU 20,000 records none 1,115 KiB 13.66 ms ~6 ms 5.94 ms 20,000 records gzip 303 KiB 31.61 ms ~23 ms 7.97 ms 20,000 records deflate 303 KiB 33.13 ms ~24 ms 8.29 ms 1 MiB random bytes none 1,024 KiB 10.13 ms ~3 ms 5.56 ms 1 MiB random bytes gzip 1,024 KiB 29.43 ms ~24 ms 4.75 ms 1 MiB random bytes gzip, disable_next_message_compression 1,024 KiB 5.98 ms ~1 ms 3.75 ms grpcio 1.84.0; server CPU from /proc in 10 ms ticks, so approximate.

2. Enable compression where it helps

Compression can be set as the server's default, the channel's default for requests, or per call:

server = grpc.aio.server(compression=grpc.Compression.Gzip)      # responses

channel = grpc.aio.insecure_channel(
    "svc:50051", compression=grpc.Compression.Gzip,                # requests
)

reply = await stub.Get(req, compression=grpc.Compression.Gzip)    # this call's request only

A handler can exempt one response from the server's default, and this is where measurement mattered. In grpcio 1.84.0's asyncio server, context.set_compression(...) had no measurable effect in either direction: with the server defaulting to gzip, setting NoCompression for the random payload left server CPU at about 23 ms per call; with the server defaulting to none, setting Gzip for the records left them at 1,115 KiB on the wire. context.disable_next_message_compression() did work:

class Recs(pbg.RecsServicer):
    async def Get(self, request, context):
        if request.random:
            context.disable_next_message_compression()    # already incompressible
        return await build_batch(request)

Measured with the server defaulting to gzip: the random response dropped to about 1 ms of server CPU and 5.98 ms per call, while the records response stayed compressed at 303 KiB. The practical policy is therefore a compressing server default with explicit exemptions — or separate servers or ports for compressible and incompressible services — verified by measuring wire bytes, not by trusting the API call. gzip and deflate produced the same 303 KiB and similar CPU here; gzip is the more widely supported choice across gRPC implementations.

Verify: a byte count on the wire confirms that compressible responses are compressed and exempted ones are not.

3. Weigh CPU against bandwidth

On a local link, compression made the call slower — 31.6 ms against 13.7 ms — because the network cost almost nothing and compression cost 17 ms of server CPU. On a constrained link, the arithmetic changes. Saving 812 KiB per response is worth about 66 ms at 100 Mbit/s and about 6.6 ms at 1 Gbit/s — a calculation from the measured sizes, not a measurement:

def transfer_ms(kib: float, mbit_per_s: float) -> float:
    return kib * 1024 * 8 / (mbit_per_s * 1_000_000) * 1000

saved = transfer_ms(1115, 100) - transfer_ms(303, 100)     # ≈ 66.5 ms at 100 Mbit/s
cost = 23 - 6                                              # extra server CPU, ms

Compression pays when the bandwidth saved exceeds the CPU spent: across regions, to mobile clients, over metered links, or when egress is billed by the byte. Within a data centre on 10 Gbit/s links, the same response spends more CPU than it saves time. Server CPU is also a capacity cost: at 23 ms per call, one core handles about 43 such responses per second, against about 160 uncompressed.

Verify: the decision is made per deployment path, using measured sizes and CPU and the real link speed.

Server CPU per 1 MiB-class response 4 horizontal bars comparing records, none with the others. Server CPU per 1 MiB-class response records, none ~6 ms records, gzip (to 27% size) ~23 ms random bytes, none ~3 ms random bytes, gzip (no saving) ~24 ms Incompressible payloads cost the same CPU and save nothing.

4. Leave small messages alone

Most RPCs carry small messages, where the question is overhead rather than savings. Measured with a single-record response over 3,000 calls: 433 µs per call with no compression, 412 µs with the server defaulting to gzip, and 386–506 µs with compression switched per RPC — differences within run-to-run noise. Compressing a few hundred bytes neither helps nor hurts measurably, so the policy can focus on large responses. gRPC compresses each message separately, so in a server-streaming RPC each chunk is compressed on its own:

async def DownloadStream(self, request, context):       # server default: gzip
    async for chunk in read_chunks(request.path, size=64 * 1024):
        if chunk_is_compressed(chunk):
            context.disable_next_message_compression()    # per message, as measured above
        yield pb.Chunk(data=chunk)

Chunks of tens of KiB still compress well when their content is compressible; very small chunks compress poorly because each starts with an empty dictionary. For large transfers, combine streaming with compression, as in sending large messages with gRPC AsyncIO.

Deciding compression for a gRPC response A flow of 5 stages. Deciding compression for a gRPC response Measure wire bytes with and without gzip Small or incompressible? leave uncompressed Link slow or metered? savings beat ~17 ms CPU Enable server default gzip Exempt + confirm disable_next_message_compression, count bytes Decide per deployment path, not once for all traffic.

Verify: small RPCs are left at the server's default, and streamed chunks are large enough to compress.

5. Watch both sides after enabling it

Compression moves work: the sender spends CPU compressing, the receiver spends CPU decompressing. Measured on the client: 5.94 ms of CPU per records call uncompressed and 7.97–8.29 ms compressed — a smaller increase than the server's, since decompression is cheaper than compression. Record both:

start_cpu = time.process_time()
reply = await stub.Get(req)
rpc_cpu_ms.labels(method="Get").observe((time.process_time() - start_cpu) * 1000)

time.process_time() measures the whole process, so this is only an approximation in a busy service; per-RPC CPU is better taken from a load test that runs one RPC type at a time. After enabling compression in production, compare server CPU, p99 latency and egress bytes with the previous week; all three should move in the expected direction.

Verify: server CPU, latency and egress are compared before and after the change, and the change is reverted where CPU rose without a matching latency or cost benefit.

Verification

gRPC compression is configured well when:

  • Wire sizes are measured for large responses, with and without compression.
  • Incompressible responses are exempted with disable_next_message_compression(), confirmed by wire bytes.
  • The CPU cost is justified by bandwidth or egress savings on the actual link.
  • Small messages are left alone, since they showed no measurable difference.

Diagnostic Hook: when a gRPC server's CPU rises after enabling compression with no drop in egress, look at what the large responses contain. A 1 MiB random payload cost about 24 ms of compression CPU per call and left the wire size at 1,024 KiB.

Pitfalls & edge cases

  • Compressing already-compressed data. Measured: no saving, about 24 ms of CPU per call.
  • Judging compression on a local link. It made calls slower: 31.6 ms against 13.7 ms.
  • Trusting context.set_compression. It changed nothing measurable in grpcio 1.84.0's aio server.
  • Tiny streamed chunks. Each chunk is compressed on its own.

Frequently Asked Questions

How do I enable gzip compression in grpc.aio?

Pass compression=grpc.Compression.Gzip to grpc.aio.server for responses or to the channel for requests. To exempt one response, call context.disable_next_message_compression(); set_compression had no effect here.

How much does gRPC gzip compression save?

For a response of 20,000 user records, 1,115 KiB became 303 KiB. Random bytes stayed at 1,024 KiB. Measure your own payloads.

Does gRPC compression slow down calls?

On a fast link, yes: server CPU rose from about 6 to 23 ms per 1 MB response and the call took 31.6 ms instead of 13.7 ms locally. It pays on slow or metered links.

Is deflate better than gzip for gRPC?

They gave the same 303 KiB and similar CPU here. Gzip is the more widely supported choice across gRPC implementations.