Measuring Memory per Connection in asyncio Servers¶
The capacity of a server holding many long-lived connections — WebSockets, streaming APIs, chat, device gateways — is usually set by memory per connection, not CPU. The number varies by an order of magnitude with the protocol stack, so measure it for the stack you run. Measured on Python 3.14 by opening 10,000 idle connections from a separate process and reading the server's resident memory before and after: a plain asyncio streams server used 6,102 bytes per connection; the same over TLS used 306,026 bytes — asyncio's TLS layer allocates a 256 KiB read buffer for every connection. WebSocket servers used 15,796 bytes with websockets, 16,101 with aiohttp and 26,407 with uvicorn and Starlette; enabling permessage-deflate raised websockets to 53,472. Shrinking asyncio's TLS buffer to 16 KiB cut TLS to 60,241 bytes per connection at the cost of bulk throughput on one connection, 1,103–1,111 instead of 1,655–1,710 MiB/s; uvloop's TLS used 45,740. This guide measures your stack the same way and turns the result into a capacity number.
Prerequisites¶
- Python 3.11+ on Linux, and a file-descriptor limit above the connection count (
ulimit -n). - Memory profiling basics, from finding memory leaks in asyncio with tracemalloc.
- The topic overview, Memory & Resource Leaks.
1. Measure resident memory from outside¶
Run the server in one process and the clients in another, so the client's memory does not pollute the number. Read the server's resident set size from /proc before and after the connections open:
python server.py & SERVER=$!
sleep 2; BEFORE=$(awk '/VmRSS/{print $2}' /proc/$SERVER/status)
python clients.py 10000 & # opens 10,000 connections and holds them
sleep 10; AFTER=$(awk '/VmRSS/{print $2}' /proc/$SERVER/status)
echo "$(( (AFTER - BEFORE) * 1024 / 10000 )) bytes per connection"
Resident memory includes everything the process touched — Python objects, buffers inside C extensions, OpenSSL state — which tracemalloc alone does not see. Each client sends one small message so that per-connection objects that are created lazily on first data exist. Measured for a plain asyncio.start_server handler that reads in a loop: 6,102 bytes per connection, and 10,007 file descriptors open in the server.
Verify: the server's file-descriptor count rises by the number of connections, confirming they are all open when memory is read.
2. Find the dominant term¶
The large jumps have specific causes. For TLS, asyncio's SSLProtocol allocates a read buffer of max_size, 256 KiB, per connection when it is created — measured, 306,026 bytes per connection in total, of which the buffer is 262,144. With 10,000 idle TLS connections, the server's resident memory grew by about 2.9 GiB. For WebSockets with permessage-deflate, each connection holds a zlib compression and decompression context; measured, the per-connection cost rose from 15,796 to 53,472 bytes when compression was negotiated:
from websockets.asyncio.server import serve
async with serve(handler, "0.0.0.0", 8765, compression=None): # 15.8 KB per connection
...
async with serve(handler, "0.0.0.0", 8765): # deflate default: 53.5 KB
...
Whether compression is worth 38 KB per connection depends on message sizes and bandwidth costs, as compared in compressing WebSocket messages with permessage-deflate.
Verify: for the stack's largest per-connection cost, the source is identified — a buffer, a compressor, a framework object — and a measurement without it confirms the difference.
3. Reduce TLS buffer cost, knowing the trade-off¶
The TLS buffer size is a class attribute on asyncio's private SSLProtocol. Lowering it before the server starts reduces memory per connection:
import asyncio.sslproto
asyncio.sslproto.SSLProtocol.max_size = 16 * 1024 # private attribute: pin and test your Python version
Measured: 60,241 bytes per connection instead of 306,026 — a fifth. Bulk throughput on a single TLS connection fell from 1,655–1,710 MiB/s to 1,103–1,111 MiB/s, because each read moves less data. For a gateway with many mostly idle connections, that trade is usually right; for a few high-bandwidth streams, it is not. Because the attribute is private, guard the change with a test that fails when a Python upgrade removes or renames it. Two alternatives avoid patching: running on uvloop, whose TLS implementation measured 45,740 bytes per connection, or terminating TLS in a proxy in front of the Python process, which then holds plain connections at about 6 KB each.
Verify: a test asserts the patched attribute exists and the per-connection measurement reflects it.
4. Turn the measurement into capacity¶
Per-connection memory gives a ceiling on connections per process, after reserving memory for everything else:
def max_connections(container_mib: int, baseline_mib: int, bytes_per_conn: int,
headroom: float = 0.3) -> int:
usable = (container_mib - baseline_mib) * 1024 * 1024 * (1 - headroom)
return int(usable / bytes_per_conn)
max_connections(1024, 80, 26_407) # uvicorn WebSockets in 1 GiB: about 26,000
max_connections(1024, 80, 306_026) # asyncio TLS in 1 GiB: about 2,300
The headroom covers message buffers under load, which idle measurements do not include — a connection receiving a message holds that message in memory too. Enforce the ceiling: refuse or shed new connections above it rather than letting the out-of-memory killer pick a moment, as in load shedding when the event loop is overloaded.
Verify: the configured connection limit per process is derived from a measured per-connection cost and the container's memory limit.
5. Re-measure under realistic traffic¶
Idle connections give the floor. Re-run the measurement with the connections doing what production connections do — periodic messages, subscriptions, pings — and after they have been open for a while:
async def client(uri, interval=5.0):
async with connect(uri) as ws:
await ws.send(json.dumps({"subscribe": "prices"}))
while True:
await ws.recv() # receive broadcasts
await asyncio.sleep(interval)
Memory per connection under traffic is typically higher than idle — per-connection queues, receive buffers, subscription state — and if it keeps rising with time at a constant connection count, that is a leak rather than a cost; see tracking task growth in long-running services. Repeat after dependency upgrades: the figures here are for specific versions of websockets, aiohttp, uvicorn and Python.
Verify: per-connection memory under realistic traffic is measured, recorded with the library versions, and stable over an hour at a constant connection count.
Verification¶
Per-connection memory is known when:
- It is measured from outside the process, as resident memory growth over thousands of connections.
- The dominant cost is identified, such as asyncio's 256 KiB TLS buffer or a deflate context.
- The connection limit per process is derived from it, with headroom.
- Measurements are repeated under traffic and after upgrades.
Diagnostic Hook: when a TLS-serving asyncio process uses gigabytes with only a few thousand idle connections, divide its resident memory by the connection count. About 300 KB per connection points to asyncio's per-connection TLS read buffer — 306,026 bytes each in this test.
Pitfalls & edge cases¶
- Measuring with tracemalloc only. It misses buffers allocated in C.
- Assuming TLS is cheap per connection. Measured: 50 times plain TCP.
- Patching
SSLProtocol.max_sizewithout a test. It is a private attribute. - Capacity from idle measurements alone. Messages in flight need memory too.
Frequently Asked Questions¶
How much memory does an asyncio connection use?
About 6.1 KB for a plain TCP stream, 16-26 KB for a WebSocket depending on the library, and 306 KB with asyncio's TLS in this test.
Why does asyncio TLS use so much memory per connection?
asyncio's SSLProtocol allocates a 256 KiB read buffer per connection. Lowering SSLProtocol.max_size to 16 KiB cut it to 60 KB, with lower bulk throughput.
Does permessage-deflate increase WebSocket memory?
Yes: from 15,796 to 53,472 bytes per connection with websockets, for the compression contexts kept per connection.
How many WebSocket connections fit in one Python process?
Divide usable memory by measured bytes per connection: at 26 KB, about 26,000 in 1 GiB with 30% headroom.
Related¶
- Memory & Resource Leaks — up to the topic overview.
- Catching leaks in CI with memory budgets — keeping these numbers from drifting.
- Resilience, Cancellation & Error Handling — the section overview.