Implementing Idle Timeouts for Connections¶
Connections that stop sending — clients that vanished without closing, peers behind a dead NAT, abandoned keep-alive sockets — hold file descriptors, memory and handler tasks until something closes them. An idle timeout reclaims them. There are two practical designs for asyncio servers, and they trade precision for overhead. Measured with 5,000 connections that chatted for one second and then went silent, against a 2-second idle limit: wrapping each read in asyncio.timeout(2.0) closed every connection 2.00–2.04 s after its last byte, and the process spent 1.86 s of CPU during the second of chatter; a single sweeper task that checks last-activity timestamps every 0.5 s closed them 2.10–2.23 s after their last byte and spent 1.73 s — about 7% less, because no timer is created and cancelled per read. Both reclaimed all 5,000. This guide implements both and explains which to use.
Prerequisites¶
- Python 3.11+, stdlib only.
- A stream server, from writing a TCP server with asyncio.start_server.
- Read deadlines, from adding read timeouts to asyncio streams.
1. Time out each read¶
The simplest idle timeout bounds every wait for data. When nothing arrives within the limit, the connection is idle and is closed:
IDLE = 300.0
async def handle(reader: asyncio.StreamReader, writer: asyncio.StreamWriter) -> None:
try:
while True:
try:
async with asyncio.timeout(IDLE):
data = await reader.read(65536)
except TimeoutError:
log.info("closing idle connection %s", writer.get_extra_info("peername"))
break
if not data:
break
await process(data, writer)
finally:
writer.close()
Measured: connections closed 2.00–2.04 s after their last byte with a 2-second limit — precise to the timer resolution. Each read creates and cancels a timer handle, which is cheap but not free: 1.86 s of CPU during a second in which 5,000 connections each sent ten messages, against 1.73 s for the sweeper. For most servers that difference is noise; the per-read timeout is the right default because it is local, obvious and exact.
Verify: a test client that connects, sends once and goes silent is disconnected after the limit.
2. Use a sweeper for very many connections¶
For tens of thousands of mostly idle connections, a single task that periodically closes connections whose last activity is too old avoids per-read timers entirely:
class IdleSweeper:
def __init__(self, idle: float, period: float = 1.0) -> None:
self.idle, self.period = idle, period
self.last_seen: dict[asyncio.StreamWriter, float] = {}
def touch(self, writer: asyncio.StreamWriter) -> None:
self.last_seen[writer] = time.monotonic()
def forget(self, writer: asyncio.StreamWriter) -> None:
self.last_seen.pop(writer, None)
async def run(self) -> None:
while True:
await asyncio.sleep(self.period)
cutoff = time.monotonic() - self.idle
for writer, seen in list(self.last_seen.items()):
if seen < cutoff:
self.forget(writer)
writer.transport.abort() # wakes the handler's read with an error
async def handle(reader, writer, sweeper: IdleSweeper) -> None:
sweeper.touch(writer)
try:
while data := await reader.read(65536):
sweeper.touch(writer)
await process(data, writer)
except ConnectionError:
pass
finally:
sweeper.forget(writer)
writer.close()
Measured: all 5,000 connections closed between 2.10 and 2.23 s of idleness — the limit plus up to one sweep period — and CPU during chatter was 7% lower. The sweep's own cost grows with the number of connections per pass, so for very large numbers, bucket connections by expiry time (a timing wheel) instead of scanning all of them. Aborting the transport makes the handler's pending read fail, so its finally runs and the per-connection state is released in one place.
Verify: with the sweeper running, the number of open connections falls back to the active ones within one idle period plus one sweep period.
3. Decide what counts as activity¶
"Idle" needs a definition. Reads alone are the usual choice, but some protocols need more:
async def write_with_activity(writer, sweeper, data: bytes) -> None:
writer.write(data)
await writer.drain()
sweeper.touch(writer) # server-push protocols: sending counts as activity too
async def on_ping(writer, sweeper) -> None:
sweeper.touch(writer) # application-level heartbeats keep the connection alive
For request-response protocols, counting only incoming data is right: a client that stops sending is idle even if the server is still writing a long response — bound that response with a write timeout instead. For server-push protocols (subscriptions, streaming), a connection may legitimately receive data for hours without sending any; count outgoing data, or require periodic application pings from the client, as WebSocket servers do with ping/pong, covered in tuning WebSocket ping/pong heartbeats.
Verify: a long-running server-push connection with no incoming data is not closed, while a silent request-response connection is.
4. Coordinate with the other timeouts on the path¶
The server's idle timeout must fit with the idle timeouts of clients and load balancers in front of it:
# Ordering that avoids races (see the stale-connections guide):
# client keepalive expiry < server idle timeout < load balancer idle timeout
CLIENT_KEEPALIVE_EXPIRY = 60
SERVER_IDLE_TIMEOUT = 75 # e.g. Uvicorn --timeout-keep-alive, nginx keepalive_timeout
LB_IDLE_TIMEOUT = 350 # e.g. AWS NLB default
If the server closes idle connections before its clients expire them, clients occasionally send a request on a connection the server is closing at that instant — the race measured at about 1 in 60 attempts in handling stale pooled connections after idle timeouts. If a load balancer drops idle connections before the server does, the server holds connections that are already dead and only notices on the next write. Write the three values down together, in one place, for every service.
Verify: configuration review shows client expiry < server idle timeout < load balancer timeout for each service pair.
5. Measure idle closes and tune the limit¶
The right limit depends on traffic: too short and active-but-quiet clients reconnect constantly; too long and dead connections accumulate. Measure both effects:
IDLE_CLOSES = Counter("connections_idle_closed_total", "Connections closed by idle timeout")
CONNECTION_AGE = Histogram("connection_age_seconds", "Connection lifetime at close",
buckets=(1, 10, 60, 300, 900, 3600, 14400))
def on_close(writer, reason: str, opened_at: float) -> None:
CONNECTION_AGE.observe(time.monotonic() - opened_at)
if reason == "idle":
IDLE_CLOSES.inc()
Compare idle closes with new connections: if most new connections follow an idle close from the same client, the limit is shorter than the client's natural pauses and is causing reconnect churn. If open connections grow while traffic is flat, dead connections are not being reclaimed — the limit is too long, or activity is counted from the wrong direction. Linux's TCP keepalive can complement both by detecting peers that disappeared without closing, as in tuning TCP keepalive for long-lived async connections.
Verify: open connection count tracks active clients, and reconnects caused by idle closes are a small share of new connections.
Verification¶
Idle timeouts work when:
- Every connection has a bound on silence, per read or via a sweeper.
- Activity is defined per protocol, including writes or pings where needed.
- Timeouts are ordered along the client, server and load balancer path.
- Idle closes and connection ages are measured and the limit tuned from them.
Diagnostic Hook: chart open connections against active connections (those with traffic in the last minute). A widening gap is connections nobody uses that the idle timeout is not reclaiming — check that the handler counts activity correctly and that the sweeper, if used, is still running.
Pitfalls & edge cases¶
- No idle timeout. Abandoned connections accumulate until descriptors run out.
- Server timeout shorter than client keep-alive. Requests race the close.
- Counting only reads on push protocols. Healthy subscribers get disconnected.
- Sweeping huge maps every second. Use a timing wheel past tens of thousands.
Frequently Asked Questions¶
How do I close idle connections in an asyncio server?
Wrap each read in async with asyncio.timeout(idle_seconds) and close the connection when TimeoutError fires. In testing this closed 5,000 idle connections within 2.00-2.04 s of a 2-second limit.
Is a timeout per read expensive with many connections?
Slightly: in testing it used about 7% more CPU than a single sweeper task while 5,000 connections were active. For typical servers the difference is negligible.
How do I implement an idle timeout without a timer per connection?
Record each connection's last activity time and run one task that periodically aborts connections whose timestamp is older than the limit. Precision is the limit plus one sweep period.
What idle timeout should a server use?
Longer than the clients' keep-alive expiry and shorter than any load balancer's idle timeout in front of it, tuned so idle closes do not cause frequent reconnects.
Related¶
- Timeouts & Deadlines — up to the topic overview.
- Enforcing request timeouts in ASGI servers — bounding active requests rather than idle connections.
- Resilience, Cancellation & Error Handling — the section overview.