Resuming WebSocket Streams After Reconnect¶
Reconnecting a WebSocket client is half of surviving a dropped connection; the other half is getting the events that were sent while it was gone. That needs sequence numbers, a replay buffer on the server, and careful ordering between replay and live delivery. Measured on Python 3.14 with websockets 17.1, a server emitting 200 events per second, and a client that dropped its connection four times for 1 second each: reconnecting without resume lost 788 events. Sending the last sequence number on reconnect and replaying from a buffer of 1,000 events lost 0, with 0 duplicates. With a buffer of only 100 events — half a second — every outage was too long to replay, and the server sent a snapshot instead, 4 times. And two ordering mistakes in the replay code each lost 316 events when replay was paced like a real network: subscribing to live events only after replaying, and advancing the client's cursor to the buffer's newest event rather than the last one actually sent. This guide builds resume correctly.
Prerequisites¶
- websockets, and a client that reconnects, from reconnecting WebSocket clients with backoff.
- Fan-out to many clients, from broadcasting to thousands of WebSocket clients.
- The topic overview, WebSocket & Real-Time Streams.
1. Number every event¶
The server gives each event a monotonically increasing sequence number and keeps the most recent ones in a bounded buffer:
class Feed:
def __init__(self, keep: int = 1000):
self.seq = 0
self.buffer: deque[dict] = deque(maxlen=keep)
self.subscribers: set[asyncio.Queue] = set()
def publish(self, payload: dict):
self.seq += 1
event = {"seq": self.seq, **payload}
self.buffer.append(event)
for q in self.subscribers:
q.put_nowait(event)
Measured without resume — a client that reconnected and simply took whatever was live — four 1-second outages at 200 events per second lost 788 events, about 200 per outage. The client could detect each gap from the jump in sequence numbers, but could not fill it. With more than one server process, the sequence must come from a shared source, such as a Redis stream ID or a database sequence, as in scaling WebSockets across processes with Redis pub/sub.
Verify: every event a client receives carries a sequence number, and a client-side check counts gaps.
2. Send the last sequence on reconnect¶
The client remembers the last sequence number it processed and sends it as the first message after connecting:
async def consume(uri: str, handle):
last_seq = None
async for ws in connect(uri): # websockets reconnects with backoff
try:
await ws.send(json.dumps({"last_seq": last_seq}))
async for raw in ws:
msg = json.loads(raw)
if msg.get("type") == "snapshot":
await load_snapshot(msg)
last_seq = msg["seq"]
continue
if last_seq is not None and msg["seq"] <= last_seq:
continue # duplicate after replay: skip
await handle(msg)
last_seq = msg["seq"]
except websockets.ConnectionClosed:
continue
Update last_seq only after an event is handled, so an event received but not processed is requested again. Measured with a 1,000-event buffer: 0 lost and 0 duplicates across the four outages. The duplicate check costs one comparison and protects against any replay overlap.
Verify: after a forced disconnect, the client's first message carries the sequence of the last handled event.
3. Subscribe before replaying, and advance per event¶
On the server, the order of operations decides whether events are lost at the seam between replay and live delivery:
async def handler(ws):
hello = json.loads(await ws.recv())
last = hello.get("last_seq")
queue: asyncio.Queue = asyncio.Queue()
feed.subscribers.add(queue) # 1. start collecting live events
try:
if last is not None:
oldest = feed.buffer[0]["seq"] if feed.buffer else feed.seq + 1
if last + 1 < oldest:
await send_snapshot(ws) # step 4
last = feed.seq
else:
for event in list(feed.buffer): # 2. replay what was missed
if event["seq"] > last:
await ws.send(json.dumps(event))
last = event["seq"] # 3. advance to what was sent
else:
last = feed.seq
while True:
event = await queue.get() # 4. live events, skipping overlap
if event["seq"] > last:
await ws.send(json.dumps(event))
last = event["seq"]
finally:
feed.subscribers.discard(queue)
With replay paced at 2 ms per event, as a real network would pace it, two plausible variations each lost 316 events over the four reconnects. Subscribing only after the replay loop missed every event published while replay was in progress. Subscribing first but setting last to the newest event in the buffer after replay — feed.buffer[-1]["seq"] — skipped events that had entered the buffer after the replay's snapshot of it, since the live queue's copies were then discarded as old. Subscribing first and advancing last per sent event lost 0. With unpaced replay, all three versions lost nothing, which is why such bugs survive local tests.
Verify: a test that slows replay — an await asyncio.sleep per replayed event — still shows 0 lost events.
4. Fall back to a snapshot when the buffer is too short¶
A buffer covers a bounded outage: 1,000 events at 200 per second is 5 seconds. When the client's sequence is older than the oldest buffered event, replay cannot fill the gap, and the server must say so rather than silently skipping. Measured with a 100-event buffer and 1-second outages: the server detected the gap on all four reconnects and sent a snapshot each time, and the client reloaded its state from it.
async def send_snapshot(ws):
state = await build_current_state() # e.g. the current order book
await ws.send(json.dumps({"type": "snapshot", "seq": feed.seq, "state": state}))
The snapshot carries the sequence it is consistent with, so the client resumes from there. Size the buffer for typical reconnect times — a few seconds for network blips, longer for mobile clients — and let the snapshot path handle the rest. Track how often snapshots are sent: a rising rate means the buffer is too small or clients are disconnecting for longer.
Verify: a client offline for longer than the buffer receives a snapshot and then a continuous stream, with no silent gap.
5. Bound the cost of replay¶
Replay sends a burst — up to the whole buffer — to one client right after it connects, on top of live traffic. With many clients reconnecting at once after a server restart, every one of them asks for a replay. Keep the burst bounded and observable:
REPLAY_LIMIT = 1000
missed = [e for e in feed.buffer if e["seq"] > last]
if len(missed) > REPLAY_LIMIT:
await send_snapshot(ws)
else:
for event in missed:
await ws.send(json.dumps(event)) # drain() back-pressure applies per send
ws.send waits when the connection's write buffer is full, so a slow client slows only its own replay. Encode each event once and reuse the encoded form for every client that needs it, and jitter client reconnects so a restart does not produce a replay storm, as in reconnecting WebSocket clients with backoff.
Verify: a test that reconnects many clients at once keeps server memory and latency for live clients within bounds.
Verification¶
Stream resume is correct when:
- Every event has a sequence number, from a source shared by all server processes.
- Clients send the last handled sequence on every reconnect and skip duplicates.
- The server subscribes before replaying and advances its cursor per event sent.
- Outages longer than the buffer produce a snapshot, never a silent gap.
Diagnostic Hook: when clients report occasional missing events after reconnects that only happen under load, slow the replay in a test with a short sleep per event. Two ordering bugs lost 316 events that way here while losing none with an unpaced replay.
Pitfalls & edge cases¶
- Reconnecting without resume. Measured: 788 events lost in four 1 s outages.
- Subscribing after replay. Measured: 316 lost with paced replay.
- Setting the cursor to the buffer's newest event. Measured: 316 lost.
- A buffer shorter than typical outages. Every reconnect becomes a snapshot.
Frequently Asked Questions¶
How do I resume a WebSocket stream after reconnecting?
Number events, keep a replay buffer, have the client send its last handled sequence on connect, and replay newer events before switching to live delivery.
What happens if the client was offline longer than the buffer?
The server should send a snapshot with the sequence it reflects. With a 100-event buffer and 1 s outages at 200 events/s, all 4 reconnects got snapshots.
Why do I lose events right after a WebSocket reconnect?
Usually the seam between replay and live events: subscribing after replay, or advancing the cursor past what was sent. Each lost 316 events with paced replay.
How large should a WebSocket replay buffer be?
Large enough for typical outages at your event rate: 1,000 events covered 5 seconds at 200/s. Longer outages should fall back to a snapshot.
Related¶
- WebSocket & Real-Time Streams — up to the topic overview.
- Limiting WebSocket message size — keeping snapshots within the message limit.
- Network I/O & Protocol Handling — the section overview.