Handling Stale Pooled Connections After Idle Timeouts¶
A pooled connection that sat idle may no longer work when you pick it up. Two very different things can have happened to it. If the server closed it after its own idle timeout, it sent a FIN, and the client can see the socket is closed before reusing it. If a middlebox — a NAT gateway, a cloud load balancer, a firewall — dropped its state for the idle connection, nothing was sent at all; the client believes the connection is fine, writes a request into it, and waits for an answer that never comes. Tested with httpx 0.28 and aiohttp 3.14: after a server closed an idle connection, the next request succeeded in 1–2 ms in both clients — they noticed and opened a new connection. Through a proxy that silently dropped connections idle for more than 2 s, the next request after 3 s idle hung for 5.0 s until the read timeout, in both clients. Setting the client's keepalive expiry below the middlebox's idle timeout (1 s against 2 s) made the same request succeed in 1 ms. This guide covers both cases and the race between them.
Prerequisites¶
- Python 3.11+, httpx and/or aiohttp.
- Pool settings, from configuring httpx limits and pool timeouts and configuring aiohttp TCPConnector limits.
- TCP keepalive, from tuning TCP keepalive for long-lived async connections.
1. Tell the two failure modes apart¶
The symptoms differ, and so do the fixes:
# Server closed the idle connection (it sent FIN): clients detect it before reuse
async with httpx.AsyncClient(limits=httpx.Limits(keepalive_expiry=30)) as client:
await client.get(url) # server keep-alive timeout is 2 s
await asyncio.sleep(3)
await client.get(url) # measured: ok in 0.002 s, on a new connection
# A middlebox silently forgot the connection: nothing tells the client
async with httpx.AsyncClient(limits=httpx.Limits(keepalive_expiry=30),
timeout=httpx.Timeout(5.0)) as client:
await client.get(via_proxy) # proxy drops state after 2 s idle
await asyncio.sleep(3)
await client.get(via_proxy) # measured: ReadTimeout after 5.005 s
The first case is handled for you: before reusing a pooled connection, httpx and aiohttp check whether the socket has been closed by the peer and discard it if so. The second case is invisible at the socket level, so the request is written into a dead connection and the client waits for its read timeout — measured at 5.0 s in both clients, with ReadTimeout from httpx and SocketTimeoutError from aiohttp. In production that shows up as occasional requests that take exactly your read timeout, usually the first request after a quiet period.
Verify: look at latency outliers: a cluster at exactly the read timeout, after idle gaps, is the silent-drop signature.
2. Expire idle connections before anything else does¶
The reliable fix is to never reuse a connection that has been idle longer than the shortest idle timeout on its path. Set the client's expiry below it:
# Typical idle timeouts on the path, pick the smallest and stay below it:
# AWS NLB 350 s, AWS ALB 60 s (configurable), GCP LB 600 s, Azure LB 240 s,
# nginx keepalive_timeout 75 s, Uvicorn --timeout-keep-alive 5 s
client = httpx.AsyncClient(limits=httpx.Limits(keepalive_expiry=4.0)) # below Uvicorn's 5 s
session = aiohttp.ClientSession(connector=aiohttp.TCPConnector(keepalive_timeout=4.0))
Measured: with keepalive_expiry=1 against a 2 s silent-drop proxy, the request after 3 s idle succeeded in 1 ms on a fresh connection. The cost is an extra handshake for traffic with gaps longer than the expiry — a few milliseconds on a LAN, more with TLS across regions — which is far cheaper than an occasional request stalling for the full read timeout. Find the actual timeouts on your path from the provider's documentation and your own proxy configuration; defaults differ between products and are often changed.
Verify: after setting the expiry, the cluster of requests at exactly the read timeout disappears.
3. Close the race with server keep-alive longer than client expiry¶
Even when the server closes connections properly, there is a race: the client picks up a connection at the same moment the server decides to close it, and the request crosses the FIN on the wire. The client sees "server disconnected without sending a response":
# The race window: client reuse and server idle close at the same instant
# httpx: httpx.RemoteProtocolError: Server disconnected without sending a response.
# aiohttp: aiohttp.ServerDisconnectedError
# Prevention: the server must keep idle connections LONGER than the client does
# uvicorn app:app --timeout-keep-alive 75 # server side
client = httpx.AsyncClient(limits=httpx.Limits(keepalive_expiry=60)) # client side, shorter
Reproduced against a server with a 2 s keep-alive timeout, by reusing a connection after idling for a random 1.995–2.010 s: 1 request in 60 failed with RemoteProtocolError. Rare, but under steady traffic with gaps near the server's timeout, it becomes a steady trickle of errors. The rule is the same one used between load balancers and backends: whichever side closes idle connections should be the client, so the server's keep-alive timeout must be longer than the client's. When the client closes first, there is no race, because the client never sends on a connection it is about to close. Uvicorn's default --timeout-keep-alive of 5 s is shorter than most clients' expiry, which makes the race likely under bursty traffic; raise it when the server sits behind a proxy or serves other services.
Verify: count RemoteProtocolError / ServerDisconnectedError by time since the connection was last used; if they cluster at the server's keep-alive timeout, the two settings are in the wrong order.
4. Retry the safe failures once¶
Some stale-connection errors will happen anyway — a server restart, a deploy, a middlebox with an undocumented timeout. They are safe to retry for idempotent requests, because the request never reached the application:
STALE = (httpx.RemoteProtocolError, httpx.ReadError, httpx.WriteError)
async def get_with_stale_retry(client: httpx.AsyncClient, url: str) -> httpx.Response:
try:
return await client.get(url)
except STALE:
return await client.get(url) # the pool discarded the dead connection; this uses a new one
httpx's transport-level retries= option only retries connection establishment failures, not errors on reused connections, so this retry belongs in your code, as discussed in retrying httpx requests with transport retries. Retry once, and only for methods that are safe to repeat — GET, HEAD, PUT and DELETE with idempotent semantics, or POST with an idempotency key. A silent drop that ends in a read timeout is not in this list: by then you have waited the full timeout, and the fix is the expiry in step 2.
Verify: a server restart during a load test produces no user-visible errors for idempotent requests.
5. Use TCP keepalive for long-lived connections¶
For connections that must stay open while idle — database connections, WebSockets, gRPC channels, message-broker links — expiring them is not an option. TCP keepalive probes keep middlebox state alive and detect dead peers:
import socket
def enable_keepalive(sock: socket.socket, idle: int = 30, interval: int = 10, count: int = 3) -> None:
sock.setsockopt(socket.SOL_SOCKET, socket.SO_KEEPALIVE, 1)
sock.setsockopt(socket.IPPROTO_TCP, socket.TCP_KEEPIDLE, idle) # first probe after 30 s idle
sock.setsockopt(socket.IPPROTO_TCP, socket.TCP_KEEPINTVL, interval) # then every 10 s
sock.setsockopt(socket.IPPROTO_TCP, socket.TCP_KEEPCNT, count) # dead after 3 misses
The probe interval must be shorter than the middlebox's idle timeout, so the middlebox sees traffic before it forgets the connection. Linux's default first probe after 2 hours is useless for this; set it per socket. Application-level pings (WebSocket ping frames, gRPC keepalive) do the same job one layer up and work through proxies that terminate TCP. The full treatment is in tuning TCP keepalive for long-lived async connections.
Verify: with keepalive enabled, a long-lived connection through the middlebox survives idle periods longer than its timeout.
Verification¶
Stale connections are handled when:
- Client keepalive expiry is below every idle timeout on the path.
- Servers keep idle connections longer than their clients do.
- Idempotent requests retry once on disconnect errors.
- Long-lived connections send keepalives more often than middleboxes expire state.
Diagnostic Hook: histogram request latency and mark the read timeout. A spike at exactly the timeout value, concentrated on requests that followed an idle gap, is silent connection loss; lower the client's expiry. Disconnect errors clustered at the server's keep-alive timeout are the close race; lengthen the server's.
Pitfalls & edge cases¶
- Client expiry longer than a load balancer's idle timeout. Measured: a 5 s hang to the read timeout.
- Server keep-alive shorter than client expiry. Produces the close race under bursty traffic.
- Relying on transport retries. httpx's
retries=does not cover reused connections. - Retrying after a read timeout. The timeout already cost the full wait; fix the expiry instead.
Frequently Asked Questions¶
Why does the first request after an idle period hang until the read timeout?
A load balancer, NAT or firewall dropped the idle connection without telling either side. The client reuses it, the request goes nowhere, and it waits for the read timeout — 5.0 s in testing with both httpx and aiohttp.
How do I fix 'Server disconnected without sending a response' in httpx?
It is usually a race with the server's keep-alive timeout. Make the client's keepalive_expiry shorter than the server's keep-alive timeout, and retry idempotent requests once on RemoteProtocolError.
What keepalive_expiry should I set in httpx?
Lower than the shortest idle timeout on the path: the server's keep-alive timeout and any load balancer or NAT idle timeout. In testing, an expiry of 1 s against a 2 s middlebox avoided the hang completely.
Do httpx and aiohttp detect connections the server closed?
Yes. When the server closed an idle connection, both discarded it and the next request succeeded in 1-2 ms on a new connection.
Related¶
- Connection Pooling & Keep-Alive — up to the topic overview.
- Warming connection pools at startup — the other end of a connection's life.
- Network I/O & Protocol Handling — the section overview.