Skip to content

Closing WebSocket Connections on Shutdown

WebSocket connections outlive deploys. When a server process stops, every connected client has to find another server, and how the server closes decides whether that is an orderly hand-over or a stampede. Tested with the websockets library and 300 connected clients: closing the server gracefully delivered close code 1001 ("going away") to all 300; aborting the connections delivered 1006 (abnormal closure) — indistinguishable, for the client, from a network failure. When all clients reconnected immediately to a second server, connections arrived at a peak of 109 per 100 ms after a graceful close and 166 after an abrupt one, enough to overflow the new server's accept backlog and push some handshakes into one-second SYN retries. With clients waiting a random 0–2 s before reconnecting, the peak fell to 21 per 100 ms. This guide closes connections in a way clients can act on, and keeps the reconnect wave manageable.

Prerequisites

1. Close with 1001 so clients know to reconnect

A close frame with code 1001 tells the client "this server is going away", which well-written clients treat as "reconnect, probably to another instance". An aborted TCP connection gives the client 1006, which looks like a network problem:

from websockets.asyncio.server import serve


async def main() -> None:
    server = await serve(handler, "0.0.0.0", 8765)
    await shutdown_requested.wait()
    server.close()                       # sends close frames with 1001 to every connection
    await server.wait_closed()           # waits for handlers to finish

Tested with 300 clients: server.close() delivered 1001 to all of them; aborting the transports delivered 1006 to all of them. The difference matters for client logic — a 1006 might reasonably trigger error reporting, longer backoff or a "connection lost" banner, while 1001 can be handled silently. In Starlette and FastAPI, close each socket yourself during shutdown with await websocket.close(code=1001), as in handling WebSockets in FastAPI and Starlette.

Verify: during a restart, client logs show close code 1001 rather than 1006.

How the server closed, and what 300 clients saw A grid of 3 rows by 3 columns. How the server closed, and what 300 clients saw shutdown close code seen reconnect peak per 100 ms server.close() 1001 going away (300/300) 109 transports aborted 1006 abnormal (300/300) 166 server.close() + client jitter 0-2 s 1001 (300/300) 21 websockets library on Python 3.14; clients reconnected to a second local server.

2. Stop accepting before closing existing connections

Close the listener first, so no new client connects to a server that is about to send it away, then close the established connections:

async def shutdown(server, hub, grace: float = 10.0) -> None:
    server.close(close_connections=False)        # stop accepting; keep existing connections
    await hub.broadcast({"type": "server_shutdown", "reconnect_in_ms": 0})   # optional heads-up
    await asyncio.sleep(1.0)                     # let clients act on the message
    for conn in list(hub.connections):
        await conn.close(code=1001, reason="server restarting")
    try:
        async with asyncio.timeout(grace):
            await server.wait_closed()
    except TimeoutError:
        for conn in list(hub.connections):
            conn.transport.abort()               # stragglers: give up

An application-level "server shutdown" message before the close lets clients that understand it start migrating gracefully — finish an upload, flush pending messages, show a subtle "reconnecting" state. Behind a load balancer, combine this with failing readiness first, so new connections already go elsewhere while existing ones are being closed.

Verify: no new connections are accepted after shutdown starts, and every existing connection receives a 1001 close.

3. Spread the reconnect wave

All clients of a stopping server reconnect at the same moment, and they all land on the remaining servers at once. Tested: immediate reconnects arrived at up to 109 per 100 ms, overflowing the new server's accept backlog; random delay of 0–2 s on the client brought the peak to 21:

// Browser client
socket.onclose = (event) => {
  const base = event.code === 1001 ? 0 : 1000;            // going away: reconnect soon
  const jitter = Math.random() * 2000;                    // spread the wave
  setTimeout(connect, base + jitter);
};

The server can help when clients are not under your control: send the shutdown notice with a per-client reconnect_in_ms value spread across a window, or close connections in batches over a few seconds rather than all at once. Spreading costs each client up to a couple of seconds of reconnection time; a stampede costs everyone handshake failures and retries.

Verify: during a rolling deploy, the connection rate at the remaining servers stays below their accept capacity, with no SYN retries visible in client timings.

Peak reconnect rate after a server stops, 300 clients 3 horizontal bars comparing abrupt close, immediate reconnect with the others. Peak reconnect rate after a server stops, 300 clients abrupt close, immediate reconnect 166 graceful close, immediate reconnect 109 graceful close, 0-2 s jitter 21 Reconnects spread over about 1.2 s without jitter because of accept-backlog overflow and SYN retries; over 2.1 s with jitter. Jitter turns a spike into a ramp.

4. Clean up per-connection state as connections close

Each connection usually has state beyond the socket: subscriptions, presence entries, per-user queues, entries in a cross-process registry. The close path must release them even during shutdown:

async def handler(ws) -> None:
    user = await authenticate(ws)
    await presence.add(user.id, instance=INSTANCE_ID)
    await hub.join(ws)
    try:
        async for message in ws:
            await handle(user, message)
    finally:
        await hub.leave(ws)
        await asyncio.shield(presence.remove(user.id, instance=INSTANCE_ID))   # must not be skipped

During shutdown, handlers are cancelled or see their connection closed; finally runs in both cases. Shield the steps that update shared state, so a second cancellation does not leave a user shown as online on an instance that no longer exists. For crash-safety, give such entries a TTL refreshed by a heartbeat, so a killed process's entries expire on their own.

Verify: after a restart, no presence or subscription entries refer to the old instance.

5. Bound the whole sequence within the grace period

WebSocket shutdown adds steps to the pod's termination budget: notice, wait, close, wait for handlers. Bound each, and abort what remains at the end:

SHUTDOWN_NOTICE = 1.0      # time for clients to act on the notice
CLOSE_BATCHES = 5          # close in batches to spread reconnects
BATCH_INTERVAL = 0.5
HANDLER_GRACE = 5.0        # handlers finishing their finally blocks


async def close_in_batches(connections: list, batches: int, interval: float) -> None:
    for i in range(batches):
        for conn in connections[i::batches]:
            await conn.close(code=1001, reason="server restarting")
        await asyncio.sleep(interval)

Here the WebSocket part needs about 1 + 2.5 + 5 = 8.5 seconds, which must fit in terminationGracePeriodSeconds alongside the readiness drain and resource cleanup. Batching on the server spreads the reconnect wave even for clients without jitter. Anything still open when the budget runs out is aborted — those clients see 1006, which is why the earlier steps should normally finish well inside the budget.

Verify: measured WebSocket shutdown time in production is below its budget, and the share of clients seeing 1006 during deploys is near zero.

How should this service close its WebSockets? A decision on What do you control with 4 outcomes. How should this service close its WebSockets? What do you control? the server only close 1001, in batches spread server-side server and clients 1001 + client jitter peak 109 -> 21 shared per-connection state shielded finally + TTL no ghosts tight grace period budget each step, abort the rest few 1006 Tell clients the truth, and do not send them all to the same place at once.

Verification

WebSocket shutdown is graceful when:

  • The listener closes first, then connections close with 1001.
  • Reconnects are spread by client jitter or server-side batching.
  • Per-connection state is cleaned in shielded finally blocks, with TTL backstops.
  • The sequence fits the grace period, with abort only for stragglers.

Diagnostic Hook: during each deploy, count close codes sent (1001 versus aborts) and chart the connection rate at the surviving instances. A spike well above the steady connection rate is a reconnect stampede — add jitter or batching; a high share of 1006 at clients means connections are being aborted rather than closed, usually because the grace period ran out.

Pitfalls & edge cases

  • Aborting instead of closing. Tested: clients saw 1006 instead of 1001.
  • Simultaneous reconnects. Tested: up to 166 per 100 ms, overflowing the accept backlog.
  • Skipped cleanup during shutdown. Ghost presence and subscriptions remain.
  • No budget for WebSockets. The grace period expires and everything is aborted.

Frequently Asked Questions

What close code should a WebSocket server send on shutdown?

1001 ("going away"). In testing, a graceful server close delivered 1001 to all 300 clients; aborting connections delivered 1006, which clients cannot tell apart from a network failure.

How do I prevent a reconnect storm when a WebSocket server restarts?

Add random delay to client reconnects or close connections in batches from the server. Jitter of 0 to 2 s cut the peak reconnect rate from 109 to 21 per 100 ms in testing.

How do I close all WebSocket connections in FastAPI on shutdown?

Keep a registry of open connections, stop accepting new ones, and call await websocket.close(code=1001) for each during the lifespan shutdown, within a time budget.

Should a WebSocket server warn clients before closing?

It helps: an application message announcing the shutdown lets clients finish pending work and reconnect calmly, before the 1001 close.