Load Testing WebSocket Servers¶
A WebSocket server is loaded by connections, not requests: thousands of long-lived sockets, each sending occasionally and receiving pushes from the server. Load testing one means opening those connections at a controlled rate, keeping them busy, and measuring two different latencies — request/response round trips and server-push delivery — while making sure the generator is not the bottleneck. Measured on Python 3.14 with a websockets 17.1 server on one core that echoed messages and broadcast to every client twice a second, driven by an asyncio generator: with the server's default listen backlog of 100, ramping 10,000 connections over five seconds caused 15,560 listen-queue overflows and a connect p99 of 14.1 s; with a backlog of 4,096, zero overflows and a connect p99 of 0.6–1.4 s. At 5,000 connections, one generator process reported an echo median of 9.1 ms; four processes with 1,250 each reported 0.3 ms for the same server. Broadcast delivery p99 rose from 42 ms at 1,000 connections to 132 ms at 5,000 and 176–224 ms at 10,000. Ten thousand idle connections took the server from 28 MiB to 530 MiB — about 50 KB each. This guide builds the generator and reads those numbers.
Prerequisites¶
- websockets 13+ (the
websockets.asyncioAPI) or another asyncio WebSocket client. - Open-loop thinking, from building an open-loop load generator in asyncio.
- The topic overview, Load Testing & Benchmarking.
1. Write a generator that models connections¶
Each simulated client opens a connection at its scheduled time, then sends a timestamped message at a fixed rate and records two latencies: the round trip of its own messages, and the delay of server pushes, measured from a timestamp the server puts in each one:
async def client(i, start_at, stop_at, rate, stats):
await asyncio.sleep(max(0, start_at - time.monotonic()))
t = time.perf_counter()
ws = await connect(URL, open_timeout=30, ping_interval=None)
stats.connect.append(time.perf_counter() - t)
async def reader():
async for raw in ws:
msg = json.loads(raw)
if "t" in msg:
stats.push.append(time.time() - msg["t"]) # server push delay
else:
stats.rtt.append(time.perf_counter() - msg["s"]) # echo round trip
read_task = asyncio.create_task(reader())
try:
while time.monotonic() < stop_at:
await ws.send(json.dumps({"s": time.perf_counter()}))
await asyncio.sleep(1 / rate)
finally:
read_task.cancel()
await ws.close()
async def run(n, ramp, hold, rate):
t0 = time.monotonic()
stats = Stats()
await asyncio.gather(*(client(i, t0 + ramp * i / n, t0 + ramp + hold, rate, stats) for i in range(n)))
return stats
Spreading start times over the ramp controls the connection rate — 10,000 over five seconds is 2,000 per second — instead of opening everything at once, which tests only the accept path. Push delay compares clocks, so it is only meaningful when generator and server share a clock, as on one machine or with tightly synchronised hosts.
Verify: at a low connection count, connect time, round trip and push delay are all close to the network's own latency.
2. Raise the listen backlog before blaming the server¶
When connections arrive faster than the server accepts them, they wait in the kernel's listen queue. If that queue is full, the kernel drops the handshake and the client retransmits after a second, then two, then four. asyncio's create_server, and the servers built on it, default to a backlog of 100:
async with serve(handler, "0.0.0.0", 8812, backlog=4096):
await asyncio.Future()
Measured with four generator processes ramping 10,000 connections over five seconds: with the default backlog, the kernel counted 15,560 ListenOverflows during the run, the connect median was 322 ms and p99 14.1 s, and the server never held more than 9,600 connections at once because some clients were still retrying. With backlog=4096 — and net.core.somaxconn at least as large, since the kernel caps the backlog at that value — there were no overflows and connect p99 fell to 0.6–1.4 s. Watch nstat -az TcpExtListenOverflows during every connection-heavy test; a non-zero delta means the results describe the kernel's queue, not the server.
Verify: the listen-overflow counter does not change during the ramp.
3. Spread the generator across processes¶
A generator process handles every message for every connection it owns — including every broadcast. At 5,000 connections and two broadcasts a second, one process parses 10,000 pushes a second in bursts, and messages wait in its event loop before being timed:
for p in 1 2 3 4; do
taskset -c $((7 + p)) python ws_load.py 1250 5 10 1 & # 4 x 1,250 connections
done
wait
Measured against the same server with 5,000 connections: one process reported an echo median of 9.1 ms and broadcast median of 80 ms; four processes with 1,250 each reported 0.3 ms and 50 ms. The single process was not at 100% CPU — 62% on average — but its CPU use was bursty, and each broadcast burst delayed everything behind it. The rule from load testing async services with Locust applies with more force here: verify that a generator's median latency stays flat when you add another generator process. If it drops, the generator was part of what you measured.
Verify: splitting the same load over more generator processes does not change the reported latencies.
4. Measure push delay as connections grow¶
The server's cost grows with connections in two ways: echo traffic scales with total message rate, and every broadcast is one send per connection. Run the same test at several connection counts and plot both latencies:
async def broadcast_tick(clients: set):
while True:
await asyncio.sleep(0.5)
broadcast(clients, json.dumps({"t": time.time()})) # one frame per client
Measured with one generator process at 1,000 connections and four above that: at 1,000, echo p99 24–27 ms and broadcast p99 42–46 ms; at 5,000, echo p99 142–145 ms and broadcast p99 131–133 ms; at 10,000, echo p99 196–533 ms and broadcast p99 176–224 ms, with the server's single core at 46–64% averaged over the run but saturated during each broadcast. Writing 10,000 frames is one burst of work on the loop, and every echo that arrives during it waits. The numbers point at the fixes: shard connections across processes, as in scaling WebSockets across processes with Redis pub/sub, and keep per-connection work in the broadcast path to a single pre-encoded frame.
Verify: you have p99 round-trip and push latency at three or more connection counts, and know the count where p99 exceeds your target.
5. Record server memory per connection¶
Connection count is usually limited by memory before CPU. Measure the server's resident memory with no clients and with all of them connected but idle, on a freshly started server:
async def report_memory(clients: set):
while True:
await asyncio.sleep(1)
rss_kib = int(open("/proc/self/status").read().split("VmRSS:")[1].split()[0])
print(f"clients={len(clients)} rss={rss_kib / 1024:.1f} MiB", flush=True)
Measured: 27.8 MiB before any connections, 530.1 MiB with 10,000 idle connections still receiving broadcasts — about 50 KB per connection for the websockets protocol object, its buffers, the asyncio transport and a handler task. After all clients disconnected, RSS stayed at 594 MiB; on a server that had run several busy rounds it stayed near 1,025–1,078 MiB with no clients, growing slowly from round to round. Memory that Python's allocator holds after a peak is normal; memory that keeps growing at the same connection count is a leak, and the difference only shows over repeated runs — the approach in measuring memory per connection covers it in depth. Size instances from the per-connection figure at your peak, plus headroom for broadcast bursts.
Verify: memory per connection is known from a fresh server, and RSS after repeated test rounds stabilises rather than growing.
Verification¶
A WebSocket load test is meaningful when:
- Connections are ramped at a known rate, and listen overflows stay at zero.
- The generator is spread across processes until adding one more changes nothing.
- Both round-trip and push latency are reported at several connection counts.
- Server memory per connection is measured on a fresh process, and its growth across rounds is checked.
Diagnostic Hook: in production, export the number of open connections, broadcast duration (time to write one frame to every client) and event-loop lag together. When broadcast duration approaches the broadcast interval, the loop spends all its time fanning out and every other message waits — the same knee these tests found at a few thousand connections per core.
Pitfalls & edge cases¶
- The default backlog of 100. Measured: connect p99 of 14.1 s from SYN retransmits.
- One generator process for thousands of connections. Measured: a 30× inflated echo median.
- Testing only round trips. Push latency grew differently and faster.
- Reading post-test RSS as a leak. Compare repeated rounds instead.
Frequently Asked Questions¶
How do I load test a WebSocket server in Python?
Write an asyncio generator that opens connections at a controlled rate, sends timestamped messages, and records connect time, round-trip time and server-push delay. Run several generator processes; four were needed for 5,000 connections in testing.
Why do WebSocket connections take seconds to open under load?
Usually a full listen queue: with asyncio's default backlog of 100, ramping 10,000 connections caused 15,560 overflows and a 14.1 s p99; a backlog of 4,096 removed them.
How much memory does a WebSocket connection use in Python?
About 50 KB per idle connection with websockets 17.1 in testing: 10,000 connections took a server from 28 MiB to 530 MiB.
How many WebSocket connections can one asyncio process handle?
Depends on traffic: with a broadcast every 0.5 s, push p99 was 42 ms at 1,000 connections and 176 to 224 ms at 10,000 on one core.
Related¶
- Load Testing & Benchmarking — up to the topic overview.
- Capacity planning with Little's law — turning these numbers into instance counts.
- Resilience, Cancellation & Error Handling — the section overview.