Skip to content

Compressing WebSocket Messages with permessage-deflate

WebSocket compression (the permessage-deflate extension) is negotiated during the handshake and then compresses each message with zlib, keeping the compression context between messages so repeated structure gets cheaper over time. Most servers enable it by default — the websockets library and Uvicorn both do — so the question is usually whether to keep it, and with what settings. Measured with websockets 17.1 on 1,000 distinct 1.1 KB messages: realistic JSON market ticks used 155 KiB on the wire instead of 1,092 KiB (14%); random hex strings used 58%. The costs: CPU for sending and receiving 5,000 messages rose from 9 µs to 21 µs per message (both ends together), and memory for 600 open connections rose from 37 KiB to 130 KiB per connection pair once each had exchanged a message. This guide decides when compression pays and tunes it where it does.

Prerequisites

  • Python 3.11+, pip install websockets; measured with websockets 17.1.
  • Connection memory budgets, from handling WebSockets in FastAPI and Starlette.
  • Your real message samples, because the ratio depends entirely on content.

1. Measure the ratio on your own messages

Compression helps exactly as much as your messages repeat themselves. Measure on real traffic before deciding:

import zlib


def deflate_ratio(messages: list[bytes], wbits: int = -12) -> float:
    """Approximate permessage-deflate with context takeover: one stream, flushed per message."""
    comp = zlib.compressobj(level=6, wbits=wbits, memLevel=5)
    raw = sent = 0
    for m in messages:
        out = comp.compress(m) + comp.flush(zlib.Z_SYNC_FLUSH)
        raw += len(m)
        sent += len(out) - 4                 # the 0x00 0x00 0xff 0xff tail is not sent
    return sent / raw


sample = [line.encode() for line in open("ws-sample.jsonl")]
print(f"{deflate_ratio(sample):.0%} of original size")

Measured through a counting proxy, JSON ticks with repeated keys and similar values compressed to 14%; random hex — whose only redundancy is the 16-symbol alphabet — to 58%. Already-compressed payloads (images, protobuf with compressed fields, encrypted blobs) do not shrink at all and only cost CPU. Small messages gain less than the ratio suggests in absolute bytes, because frame headers and TCP/IP overhead are unchanged.

Verify: the ratio measured on a day of sampled production messages, not on a synthetic example.

Bytes on the wire for 1,000 messages of 1.1 KB 4 horizontal bars comparing JSON ticks, uncompressed with the others. Bytes on the wire for 1,000 messages of 1.1 KB JSON ticks, uncompressed 1,092 KiB JSON ticks, deflate 155 KiB (14%) random hex, uncompressed 1,088 KiB random hex, deflate 634 KiB (58%) websockets 17.1 default compression settings; every message distinct; counted server-to-client bytes through a TCP proxy. Structured JSON shrinks dramatically; high-entropy data barely does.

2. Weigh the CPU and memory costs

Compression is paid per message in CPU and per connection in memory:

from websockets.asyncio.server import serve

# Default: permessage-deflate negotiated if the client offers it
async with serve(handler, "0.0.0.0", 8765):                       # compression="deflate"
    ...

# Disabled: no extension negotiated, no zlib contexts allocated
async with serve(handler, "0.0.0.0", 8765, compression=None):
    ...

Measured: sending and receiving 5,000 JSON messages took 9 µs of CPU per message without compression and 21 µs with it, both ends in one process. Memory for 600 connections that had each exchanged a message was 37 KiB per connection pair without compression and 130 KiB with it — each side holds a compressor and a decompressor for the life of the connection. For 50,000 connections per server, that difference is several gigabytes. Bandwidth saved must be worth that: it usually is for mobile clients and metered egress, and usually is not inside a data centre.

Verify: a soak test at target connection count stays within the memory budget with compression on, or compression is turned off.

What compression cost per message and per connection A grid of 3 rows by 3 columns. What compression cost per message and per connection measure no compression permessage-deflate CPU per message (send + receive) 9 µs 21 µs memory per connection pair 37 KiB 130 KiB JSON bytes on the wire 100% 14% Bandwidth is bought with CPU and long-lived per-connection memory.

3. Tune window size and memory level

The zlib window and memory level set the memory per context and, to a lesser degree, the ratio. The websockets library already uses conservative defaults; make them explicit if you change anything:

from websockets.extensions.permessage_deflate import ServerPerMessageDeflateFactory

compression = ServerPerMessageDeflateFactory(
    server_max_window_bits=12,          # 4 KiB history instead of 32 KiB
    client_max_window_bits=12,
    compress_settings={"memLevel": 5},  # smaller internal state
)

async with serve(handler, "0.0.0.0", 8765, compression=None, extensions=[compression]):
    ...

A smaller window remembers less history, so it finds fewer long repeats, but for messages of a few kilobytes with repeated keys the loss is small. server_no_context_takeover resets the context after every message, which frees nothing in zlib's allocation but makes each message compress independently — useful when messages share nothing, harmful when they share structure. Measure the ratio and memory with each candidate setting on your sample.

Verify: the chosen settings keep the ratio within a few percent of the default while meeting the memory budget.

4. Skip compression for small or incompressible messages

Compressing a 60-byte heartbeat or a JPEG costs CPU and saves nothing. Some servers let you send individual messages uncompressed; otherwise, shape traffic so compression applies where it helps:

# Option A: batch small updates into one larger message per interval
async def batched_sender(ws, updates: asyncio.Queue, interval: float = 0.05) -> None:
    while True:
        batch = [await updates.get()]
        await asyncio.sleep(interval)
        while not updates.empty():
            batch.append(updates.get_nowait())
        await ws.send(json.dumps(batch))          # one larger, more compressible frame


# Option B: send binary blobs on a separate, uncompressed connection
async with serve(media_handler, "0.0.0.0", 8766, compression=None):
    ...

Batching improves both sides: fewer frames, and more repeated structure per frame for deflate to exploit. Sending already-compressed media through a separate endpoint without the extension keeps the CPU for messages that benefit. A threshold below which messages are not compressed is supported by some servers; websockets compresses every message once the extension is negotiated.

Verify: CPU spent in zlib (from a profiler) is concentrated on message types with good ratios.

5. Mind proxies, browsers and security

Compression is negotiated end to end only if every hop passes the Sec-WebSocket-Extensions header. A few practical points:

async def handler(ws):
    negotiated = ws.response.headers.get("Sec-WebSocket-Extensions", "")
    log.debug("extensions: %s", negotiated or "none")       # confirms what was agreed

Browsers offer permessage-deflate automatically and cannot be told not to; the server decides by accepting or declining. Reverse proxies that terminate WebSockets (some CDNs, older load balancers) may strip the header, silently disabling compression — check the negotiated value in production. And compression of secrets alongside attacker-influenced data in the same context can leak information through message sizes (the CRIME/BREACH family of attacks); avoid sending tokens or other secrets in messages that also echo user input, or disable context takeover for those connections.

Verify: the negotiated extension is logged for a sample of production connections and matches what you configured.

Should this WebSocket use permessage-deflate? A decision on What do the messages and clients look like with 4 outcomes. Should this WebSocket use permessage-deflate? What do the messages and clients look like? repetitive JSON, mobile or metered keep deflate 14% of the bytes huge connection counts, fast network compression=None save ~93 KiB/pair binary or pre-compressed separate uncompressed endpoint no gain secrets + echoed user input no context takeover or off size side channel Decide with your own message sample and connection budget.

Verification

Compression is configured deliberately when:

  • The ratio was measured on real messages and justifies the cost.
  • Memory per connection with compression fits the connection budget.
  • Small or incompressible messages are batched or sent uncompressed.
  • The negotiated extension is confirmed through every proxy.

Diagnostic Hook: export bytes sent before and after compression per endpoint, and process memory per open connection. A ratio near 100% means compression is pure cost on that endpoint; memory per connection that jumps when compression is enabled tells you exactly what turning it off would buy.

Pitfalls & edge cases

  • Assuming compression is free. CPU per message more than doubled in testing.
  • Ignoring per-connection memory. About 93 KiB more per connection pair.
  • Compressing incompressible data. Random data still used 58% of its size.
  • Proxies stripping the extension. Compression silently off in production.

Frequently Asked Questions

Should I enable permessage-deflate for WebSockets?

For repetitive JSON to bandwidth-constrained clients, usually yes: JSON ticks shrank to 14% in testing. For very many connections on fast networks, the extra memory and CPU often outweigh the savings.

How much memory does WebSocket compression use?

In testing with websockets 17.1 defaults, about 93 KiB more per connection pair (both ends) after the first message, because each side keeps zlib compression and decompression contexts for the connection's lifetime.

How do I disable WebSocket compression in Python?

Pass compression=None to websockets' serve or connect. In Uvicorn, use --ws-per-message-deflate false.

Does permessage-deflate help with binary data?

Only if the data has redundancy. Already-compressed formats such as images do not shrink, and random-looking data compressed only to 58% in testing.