Closing WebSocket Connections on Shutdown¶
WebSocket connections outlive deploys. When a server process stops, every connected client has to find another server, and how the server closes decides whether that is an orderly hand-over or a stampede. Tested with the websockets library and 300 connected clients: closing the server gracefully delivered close code 1001 ("going away") to all 300; aborting the connections delivered 1006 (abnormal closure) — indistinguishable, for the client, from a network failure. When all clients reconnected immediately to a second server, connections arrived at a peak of 109 per 100 ms after a graceful close and 166 after an abrupt one, enough to overflow the new server's accept backlog and push some handshakes into one-second SYN retries. With clients waiting a random 0–2 s before reconnecting, the peak fell to 21 per 100 ms. This guide closes connections in a way clients can act on, and keeps the reconnect wave manageable.
Prerequisites¶
- Python 3.11+,
pip install websockets; patterns apply to Starlette and FastAPI too. - Pod shutdown sequencing, from shutting down asyncio pods in Kubernetes.
- Client reconnection, from reconnecting WebSocket clients with backoff.
1. Close with 1001 so clients know to reconnect¶
A close frame with code 1001 tells the client "this server is going away", which well-written clients treat as "reconnect, probably to another instance". An aborted TCP connection gives the client 1006, which looks like a network problem:
from websockets.asyncio.server import serve
async def main() -> None:
server = await serve(handler, "0.0.0.0", 8765)
await shutdown_requested.wait()
server.close() # sends close frames with 1001 to every connection
await server.wait_closed() # waits for handlers to finish
Tested with 300 clients: server.close() delivered 1001 to all of them; aborting the transports delivered 1006 to all of them. The difference matters for client logic — a 1006 might reasonably trigger error reporting, longer backoff or a "connection lost" banner, while 1001 can be handled silently. In Starlette and FastAPI, close each socket yourself during shutdown with await websocket.close(code=1001), as in handling WebSockets in FastAPI and Starlette.
Verify: during a restart, client logs show close code 1001 rather than 1006.
2. Stop accepting before closing existing connections¶
Close the listener first, so no new client connects to a server that is about to send it away, then close the established connections:
async def shutdown(server, hub, grace: float = 10.0) -> None:
server.close(close_connections=False) # stop accepting; keep existing connections
await hub.broadcast({"type": "server_shutdown", "reconnect_in_ms": 0}) # optional heads-up
await asyncio.sleep(1.0) # let clients act on the message
for conn in list(hub.connections):
await conn.close(code=1001, reason="server restarting")
try:
async with asyncio.timeout(grace):
await server.wait_closed()
except TimeoutError:
for conn in list(hub.connections):
conn.transport.abort() # stragglers: give up
An application-level "server shutdown" message before the close lets clients that understand it start migrating gracefully — finish an upload, flush pending messages, show a subtle "reconnecting" state. Behind a load balancer, combine this with failing readiness first, so new connections already go elsewhere while existing ones are being closed.
Verify: no new connections are accepted after shutdown starts, and every existing connection receives a 1001 close.
3. Spread the reconnect wave¶
All clients of a stopping server reconnect at the same moment, and they all land on the remaining servers at once. Tested: immediate reconnects arrived at up to 109 per 100 ms, overflowing the new server's accept backlog; random delay of 0–2 s on the client brought the peak to 21:
// Browser client
socket.onclose = (event) => {
const base = event.code === 1001 ? 0 : 1000; // going away: reconnect soon
const jitter = Math.random() * 2000; // spread the wave
setTimeout(connect, base + jitter);
};
The server can help when clients are not under your control: send the shutdown notice with a per-client reconnect_in_ms value spread across a window, or close connections in batches over a few seconds rather than all at once. Spreading costs each client up to a couple of seconds of reconnection time; a stampede costs everyone handshake failures and retries.
Verify: during a rolling deploy, the connection rate at the remaining servers stays below their accept capacity, with no SYN retries visible in client timings.
4. Clean up per-connection state as connections close¶
Each connection usually has state beyond the socket: subscriptions, presence entries, per-user queues, entries in a cross-process registry. The close path must release them even during shutdown:
async def handler(ws) -> None:
user = await authenticate(ws)
await presence.add(user.id, instance=INSTANCE_ID)
await hub.join(ws)
try:
async for message in ws:
await handle(user, message)
finally:
await hub.leave(ws)
await asyncio.shield(presence.remove(user.id, instance=INSTANCE_ID)) # must not be skipped
During shutdown, handlers are cancelled or see their connection closed; finally runs in both cases. Shield the steps that update shared state, so a second cancellation does not leave a user shown as online on an instance that no longer exists. For crash-safety, give such entries a TTL refreshed by a heartbeat, so a killed process's entries expire on their own.
Verify: after a restart, no presence or subscription entries refer to the old instance.
5. Bound the whole sequence within the grace period¶
WebSocket shutdown adds steps to the pod's termination budget: notice, wait, close, wait for handlers. Bound each, and abort what remains at the end:
SHUTDOWN_NOTICE = 1.0 # time for clients to act on the notice
CLOSE_BATCHES = 5 # close in batches to spread reconnects
BATCH_INTERVAL = 0.5
HANDLER_GRACE = 5.0 # handlers finishing their finally blocks
async def close_in_batches(connections: list, batches: int, interval: float) -> None:
for i in range(batches):
for conn in connections[i::batches]:
await conn.close(code=1001, reason="server restarting")
await asyncio.sleep(interval)
Here the WebSocket part needs about 1 + 2.5 + 5 = 8.5 seconds, which must fit in terminationGracePeriodSeconds alongside the readiness drain and resource cleanup. Batching on the server spreads the reconnect wave even for clients without jitter. Anything still open when the budget runs out is aborted — those clients see 1006, which is why the earlier steps should normally finish well inside the budget.
Verify: measured WebSocket shutdown time in production is below its budget, and the share of clients seeing 1006 during deploys is near zero.
Verification¶
WebSocket shutdown is graceful when:
- The listener closes first, then connections close with 1001.
- Reconnects are spread by client jitter or server-side batching.
- Per-connection state is cleaned in shielded
finallyblocks, with TTL backstops. - The sequence fits the grace period, with abort only for stragglers.
Diagnostic Hook: during each deploy, count close codes sent (1001 versus aborts) and chart the connection rate at the surviving instances. A spike well above the steady connection rate is a reconnect stampede — add jitter or batching; a high share of 1006 at clients means connections are being aborted rather than closed, usually because the grace period ran out.
Pitfalls & edge cases¶
- Aborting instead of closing. Tested: clients saw 1006 instead of 1001.
- Simultaneous reconnects. Tested: up to 166 per 100 ms, overflowing the accept backlog.
- Skipped cleanup during shutdown. Ghost presence and subscriptions remain.
- No budget for WebSockets. The grace period expires and everything is aborted.
Frequently Asked Questions¶
What close code should a WebSocket server send on shutdown?
1001 ("going away"). In testing, a graceful server close delivered 1001 to all 300 clients; aborting connections delivered 1006, which clients cannot tell apart from a network failure.
How do I prevent a reconnect storm when a WebSocket server restarts?
Add random delay to client reconnects or close connections in batches from the server. Jitter of 0 to 2 s cut the peak reconnect rate from 109 to 21 per 100 ms in testing.
How do I close all WebSocket connections in FastAPI on shutdown?
Keep a registry of open connections, stop accepting new ones, and call await websocket.close(code=1001) for each during the lifespan shutdown, within a time budget.
Should a WebSocket server warn clients before closing?
It helps: an application message announcing the shutdown lets clients finish pending work and reconnect calmly, before the 1001 close.
Related¶
- Graceful Shutdown & Signals — up to the topic overview.
- Scaling WebSockets across processes with Redis pub/sub — the shared state that must survive instance turnover.
- Resilience, Cancellation & Error Handling — the section overview.