Skip to content

Sending Large Messages with gRPC AsyncIO

gRPC caps received messages at 4 MB by default, and the first large payload a service sends — a file, a report, a batch of embeddings — fails with RESOURCE_EXHAUSTED. Raising the limit is the quick fix; streaming the payload in chunks is the one that scales. Measured on Python 3.14 with grpcio 1.84.0 over a local channel, moving a 50 MiB payload: with default limits, an upload failed with Received message larger than max (52428805 vs. 4194304) on the server and a download failed on the client. With both limits raised to 100 MiB, a single-message upload took 0.62 s, the client peaked at 236 MiB of resident memory, and the server's event loop stalled for 249–295 ms while it handled the message. Streaming the same 50 MiB as 64 KiB chunks took 0.09 s for the upload and 0.06 s for the download, with the client at 87–95 MiB, the server at 103–111 MiB, and loop stalls of 2.3–6.5 ms where measured. This guide covers the limits, the streaming version, and how to choose a chunk size.

Prerequisites

1. Recognize the default limit

The service used for these measurements has unary and streaming versions of upload and download:

message Chunk { bytes data = 1; }
message Ack   { int64 size = 1; }
message Req   { int64 size = 1; int64 chunk = 2; }

service Blob {
  rpc Upload (Chunk) returns (Ack);
  rpc UploadStream (stream Chunk) returns (Ack);
  rpc Download (Req) returns (Chunk);
  rpc DownloadStream (Req) returns (stream Chunk);
}

Measured with default channel and server options: uploading 50 MiB raised AioRpcError with status RESOURCE_EXHAUSTED and details SERVER: Received message larger than max (52428805 vs. 4194304), after 0.41 s spent sending it. Downloading 50 MiB raised the same status with CLIENT: Received message larger than max. The limit applies to the receiver: a server's max_receive_message_length bounds uploads, a client's bounds downloads. The default send limit is unlimited in grpcio, so the sender finds out only from the error.

Verify: the largest message each RPC can carry is known, and errors with RESOURCE_EXHAUSTED and "larger than max" are traced to a size limit rather than retried.

2. Raise the limits deliberately, if at all

Both sides take channel options:

LIMIT = 100 * 1024 * 1024
opts = [("grpc.max_receive_message_length", LIMIT), ("grpc.max_send_message_length", LIMIT)]

server = grpc.aio.server(options=opts)
channel = grpc.aio.insecure_channel("svc:50051", options=opts)

Measured with 100 MiB limits: the 50 MiB upload succeeded in 0.62 s and the download in 0.61 s. The cost showed in memory and loop time. The client peaked at 236–238 MiB of resident memory, almost five times the payload, as protobuf serialization and the transport each held copies. The server peaked at 251–285 MiB, including its own 64 MiB test buffer, and a probe on its event loop measured stalls of 249–295 ms per large message: deserializing 50 MiB into a protobuf object happened on the loop. Every other RPC on that server waited that long. Raise limits for payloads that are occasionally a little over 4 MB; for anything routinely large, stream.

Verify: a raised limit is documented with the largest expected payload, and a load test with that payload shows acceptable loop stalls on the server.

Moving 50 MiB over a local gRPC channel A grid of 6 rows by 4 columns. Moving 50 MiB over a local gRPC channel method time client peak RSS server loop stall unary, default 4 MB limit RESOURCE_EXHAUSTED - - unary upload, 100 MiB limit 0.62 s 236 MiB 249-295 ms unary download, 100 MiB limit 0.61 s 238 MiB 282 ms stream upload, 64 KiB chunks 0.09 s 87 MiB not measured stream download, 64 KiB chunks 0.06 s 95 MiB 2.5 ms stream upload / download, 1 MiB chunks 0.08 / 0.09 s 90 / 123 MiB 2.3 / 6.5 ms grpcio 1.84.0, Python 3.14; the 50 MiB payload alone accounts for 50 MiB of client memory.

3. Stream large payloads in chunks

A client-streaming upload sends the payload as a sequence of small messages, and the server consumes them as they arrive:

CHUNK = 64 * 1024

async def upload(stub, data: bytes) -> int:
    async def chunks():
        view = memoryview(data)
        for offset in range(0, len(data), CHUNK):
            yield pb.Chunk(data=bytes(view[offset:offset + CHUNK]))
    return (await stub.UploadStream(chunks())).size

class Blob(pbg.BlobServicer):
    async def UploadStream(self, request_iterator, context):
        size = 0
        async for chunk in request_iterator:
            size += len(chunk.data)            # write to disk or object storage here
        return pb.Ack(size=size)

Measured: 0.09 s for 50 MiB — seven times faster than the single message — with the client at 87 MiB and the server at 111 MiB. Each chunk is deserialized separately, so the server's loop is never blocked by more than one chunk's worth of work, and flow control keeps the sender from running far ahead of the receiver. Downloads use a server-streaming RPC the same way: 0.06 s, with the client at 95 MiB and loop stalls of 2.5 ms on the server. When the payload comes from a file, read and yield it chunk by chunk rather than loading it whole, so neither side ever holds all 50 MiB.

Verify: the streaming RPC's peak memory on both sides is close to the baseline plus a few chunks, regardless of payload size.

A client-streaming upload A sequence of 6 messages between 3 participants. A client-streaming upload client HTTP/2 transport server handler open UploadStream Chunk 1..n, 64 KiB each deliver as flow control allows async for: handle one chunk end of stream Ack(size) No message ever approaches the 4 MB limit.

4. Choose a chunk size

The chunk size trades per-message overhead against per-message work. Measured for 50 MiB: 64 KiB chunks — 800 messages — took 0.09 s up and 0.06 s down; 1 MiB chunks — 50 messages — took 0.08 s up and 0.09 s down, and the client downloading them peaked at 123 MiB instead of 95, and the server's longest loop stall rose from 2.5 to 6.5 ms. Both are far better than one 50 MiB message. Between 16 KiB and 1 MiB, the choice matters less than streaming at all:

CHUNK = 64 * 1024     # 64 KiB: low memory, short loop stalls, negligible overhead at this size

Smaller chunks — a few KiB — multiply per-message costs, which are measured in microseconds of Python per message; larger chunks lengthen the server's per-message stall. Keep chunks well below the 4 MB default limit so no option changes are needed.

Verify: a sweep of two or three chunk sizes on the real network shows throughput within a few percent of each other; pick the one with lower memory.

Upload time for 50 MiB 3 horizontal bars comparing one 50 MiB message (100 MiB limit) with the others. Upload time for 50 MiB one 50 MiB message (100 MiB limit) 0.62 s stream, 1 MiB chunks 0.08 s stream, 64 KiB chunks 0.09 s Streaming was about seven times faster and used a third of the memory.

5. Handle partial transfers

A stream can fail midway — a deadline, a dropped connection, a server restart. Make uploads resumable or idempotent, so a retry does not have to resend everything or store duplicates:

async def upload_with_id(stub, upload_id: str, data: bytes, start: int = 0) -> int:
    async def chunks():
        view = memoryview(data)
        for offset in range(start, len(data), CHUNK):
            yield pb.ChunkAt(upload_id=upload_id, offset=offset,
                             data=bytes(view[offset:offset + CHUNK]))
    return (await stub.UploadAt(chunks(), timeout=60)).size     # server returns bytes stored

The server stores chunks by upload_id and offset and reports how much it has, so a client can resume from that offset after a failure. Set a deadline on every streaming call, as described in propagating gRPC deadlines and cancellation, and size it for the payload: a fixed 5-second deadline that suits small calls will cut off large transfers.

Verify: an interrupted upload can be resumed from the server's reported offset, and deadlines scale with payload size.

Verification

Large messages are handled well when:

  • Payloads that can exceed 4 MB use streaming RPCs, with chunks of tens to hundreds of KiB.
  • Raised message limits, where used, are documented and load-tested for loop stalls.
  • Neither side holds the whole payload, measured by peak memory.
  • Streams have deadlines sized to the payload and can resume after failure.

Diagnostic Hook: when a gRPC server's unrelated RPCs slow down while one client uploads, check for unary RPCs carrying large messages under a raised limit. A 50 MiB message stalled the server's loop for up to 295 ms in this test; the same data as a stream stalled it for a few milliseconds.

Pitfalls & edge cases

  • Raising limits instead of streaming. Measured: 249–295 ms loop stalls per 50 MiB message.
  • Raising only one side's limit. The receiver's limit is the one that applies.
  • Retrying RESOURCE_EXHAUSTED for size errors. It fails identically every time.
  • Fixed short deadlines on large transfers. Size the deadline to the payload.

Frequently Asked Questions

What is gRPC's default maximum message size in Python?

4 MB (4,194,304 bytes) for received messages. A 50 MiB upload failed with RESOURCE_EXHAUSTED: Received message larger than max (52428805 vs. 4194304).

How do I increase the gRPC message size limit in grpc.aio?

Pass ("grpc.max_receive_message_length", n) and ("grpc.max_send_message_length", n) in options to grpc.aio.server and the channel. The receiver's limit is the one enforced.

Should I stream large payloads in gRPC?

Yes. 50 MiB as 64 KiB chunks took 0.09 s with the client at 87 MiB, against 0.62 s and 236 MiB as one message, which also stalled the server loop for about 250 ms.

What chunk size should gRPC streaming use?

64 KiB to 1 MiB performed similarly here, 0.06-0.09 s for 50 MiB. Smaller chunks used less memory and gave shorter loop stalls.