Compressing WebSocket Messages with permessage-deflate¶
WebSocket compression (the permessage-deflate extension) is negotiated during the handshake and then compresses each message with zlib, keeping the compression context between messages so repeated structure gets cheaper over time. Most servers enable it by default — the websockets library and Uvicorn both do — so the question is usually whether to keep it, and with what settings. Measured with websockets 17.1 on 1,000 distinct 1.1 KB messages: realistic JSON market ticks used 155 KiB on the wire instead of 1,092 KiB (14%); random hex strings used 58%. The costs: CPU for sending and receiving 5,000 messages rose from 9 µs to 21 µs per message (both ends together), and memory for 600 open connections rose from 37 KiB to 130 KiB per connection pair once each had exchanged a message. This guide decides when compression pays and tunes it where it does.
Prerequisites¶
- Python 3.11+,
pip install websockets; measured with websockets 17.1. - Connection memory budgets, from handling WebSockets in FastAPI and Starlette.
- Your real message samples, because the ratio depends entirely on content.
1. Measure the ratio on your own messages¶
Compression helps exactly as much as your messages repeat themselves. Measure on real traffic before deciding:
import zlib
def deflate_ratio(messages: list[bytes], wbits: int = -12) -> float:
"""Approximate permessage-deflate with context takeover: one stream, flushed per message."""
comp = zlib.compressobj(level=6, wbits=wbits, memLevel=5)
raw = sent = 0
for m in messages:
out = comp.compress(m) + comp.flush(zlib.Z_SYNC_FLUSH)
raw += len(m)
sent += len(out) - 4 # the 0x00 0x00 0xff 0xff tail is not sent
return sent / raw
sample = [line.encode() for line in open("ws-sample.jsonl")]
print(f"{deflate_ratio(sample):.0%} of original size")
Measured through a counting proxy, JSON ticks with repeated keys and similar values compressed to 14%; random hex — whose only redundancy is the 16-symbol alphabet — to 58%. Already-compressed payloads (images, protobuf with compressed fields, encrypted blobs) do not shrink at all and only cost CPU. Small messages gain less than the ratio suggests in absolute bytes, because frame headers and TCP/IP overhead are unchanged.
Verify: the ratio measured on a day of sampled production messages, not on a synthetic example.
2. Weigh the CPU and memory costs¶
Compression is paid per message in CPU and per connection in memory:
from websockets.asyncio.server import serve
# Default: permessage-deflate negotiated if the client offers it
async with serve(handler, "0.0.0.0", 8765): # compression="deflate"
...
# Disabled: no extension negotiated, no zlib contexts allocated
async with serve(handler, "0.0.0.0", 8765, compression=None):
...
Measured: sending and receiving 5,000 JSON messages took 9 µs of CPU per message without compression and 21 µs with it, both ends in one process. Memory for 600 connections that had each exchanged a message was 37 KiB per connection pair without compression and 130 KiB with it — each side holds a compressor and a decompressor for the life of the connection. For 50,000 connections per server, that difference is several gigabytes. Bandwidth saved must be worth that: it usually is for mobile clients and metered egress, and usually is not inside a data centre.
Verify: a soak test at target connection count stays within the memory budget with compression on, or compression is turned off.
3. Tune window size and memory level¶
The zlib window and memory level set the memory per context and, to a lesser degree, the ratio. The websockets library already uses conservative defaults; make them explicit if you change anything:
from websockets.extensions.permessage_deflate import ServerPerMessageDeflateFactory
compression = ServerPerMessageDeflateFactory(
server_max_window_bits=12, # 4 KiB history instead of 32 KiB
client_max_window_bits=12,
compress_settings={"memLevel": 5}, # smaller internal state
)
async with serve(handler, "0.0.0.0", 8765, compression=None, extensions=[compression]):
...
A smaller window remembers less history, so it finds fewer long repeats, but for messages of a few kilobytes with repeated keys the loss is small. server_no_context_takeover resets the context after every message, which frees nothing in zlib's allocation but makes each message compress independently — useful when messages share nothing, harmful when they share structure. Measure the ratio and memory with each candidate setting on your sample.
Verify: the chosen settings keep the ratio within a few percent of the default while meeting the memory budget.
4. Skip compression for small or incompressible messages¶
Compressing a 60-byte heartbeat or a JPEG costs CPU and saves nothing. Some servers let you send individual messages uncompressed; otherwise, shape traffic so compression applies where it helps:
# Option A: batch small updates into one larger message per interval
async def batched_sender(ws, updates: asyncio.Queue, interval: float = 0.05) -> None:
while True:
batch = [await updates.get()]
await asyncio.sleep(interval)
while not updates.empty():
batch.append(updates.get_nowait())
await ws.send(json.dumps(batch)) # one larger, more compressible frame
# Option B: send binary blobs on a separate, uncompressed connection
async with serve(media_handler, "0.0.0.0", 8766, compression=None):
...
Batching improves both sides: fewer frames, and more repeated structure per frame for deflate to exploit. Sending already-compressed media through a separate endpoint without the extension keeps the CPU for messages that benefit. A threshold below which messages are not compressed is supported by some servers; websockets compresses every message once the extension is negotiated.
Verify: CPU spent in zlib (from a profiler) is concentrated on message types with good ratios.
5. Mind proxies, browsers and security¶
Compression is negotiated end to end only if every hop passes the Sec-WebSocket-Extensions header. A few practical points:
async def handler(ws):
negotiated = ws.response.headers.get("Sec-WebSocket-Extensions", "")
log.debug("extensions: %s", negotiated or "none") # confirms what was agreed
Browsers offer permessage-deflate automatically and cannot be told not to; the server decides by accepting or declining. Reverse proxies that terminate WebSockets (some CDNs, older load balancers) may strip the header, silently disabling compression — check the negotiated value in production. And compression of secrets alongside attacker-influenced data in the same context can leak information through message sizes (the CRIME/BREACH family of attacks); avoid sending tokens or other secrets in messages that also echo user input, or disable context takeover for those connections.
Verify: the negotiated extension is logged for a sample of production connections and matches what you configured.
Verification¶
Compression is configured deliberately when:
- The ratio was measured on real messages and justifies the cost.
- Memory per connection with compression fits the connection budget.
- Small or incompressible messages are batched or sent uncompressed.
- The negotiated extension is confirmed through every proxy.
Diagnostic Hook: export bytes sent before and after compression per endpoint, and process memory per open connection. A ratio near 100% means compression is pure cost on that endpoint; memory per connection that jumps when compression is enabled tells you exactly what turning it off would buy.
Pitfalls & edge cases¶
- Assuming compression is free. CPU per message more than doubled in testing.
- Ignoring per-connection memory. About 93 KiB more per connection pair.
- Compressing incompressible data. Random data still used 58% of its size.
- Proxies stripping the extension. Compression silently off in production.
Frequently Asked Questions¶
Should I enable permessage-deflate for WebSockets?
For repetitive JSON to bandwidth-constrained clients, usually yes: JSON ticks shrank to 14% in testing. For very many connections on fast networks, the extra memory and CPU often outweigh the savings.
How much memory does WebSocket compression use?
In testing with websockets 17.1 defaults, about 93 KiB more per connection pair (both ends) after the first message, because each side keeps zlib compression and decompression contexts for the connection's lifetime.
How do I disable WebSocket compression in Python?
Pass compression=None to websockets' serve or connect. In Uvicorn, use --ws-per-message-deflate false.
Does permessage-deflate help with binary data?
Only if the data has redundancy. Already-compressed formats such as images do not shrink, and random-looking data compressed only to 58% in testing.
Related¶
- WebSocket & Real-Time Streams — up to the topic overview.
- Broadcasting to thousands of WebSocket clients — where per-message CPU multiplies by audience size.
- Network I/O & Protocol Handling — the section overview.