Configuring httpx Limits and Pool Timeouts¶
An httpx.AsyncClient keeps a pool of connections, sized by httpx.Limits, and waits for a free one under the pool timeout. The defaults — 100 connections, 20 kept alive, 5 s for every timeout — are reasonable for many services and wrong for some. The pool timeout in particular does not mean what it seems to. Tested with httpx 0.28 against a server answering in 0.2 s, 50 concurrent requests through a pool limited to 5 connections with pool=0.5: none raised PoolTimeout, though the median request waited 1.23 s and the slowest 2.04 s — the timeout never fired because a connection became free every 0.2 s. With a 0.6 s server, 45 of 50 failed with PoolTimeout at 0.5 s. To bound how long a request may queue in total, wrap it in asyncio.timeout. This guide sets each limit deliberately and explains what each timeout does and does not bound.
Prerequisites¶
- Python 3.11+,
pip install httpx; measured with httpx 0.28.1 and httpcore 1.0.9. - Pool sizing, from sizing async connection pools for throughput.
- Timeout layers, from setting connect, read and total timeouts in async HTTP clients.
1. Set the three limits¶
httpx.Limits has three knobs: how many connections may exist, how many idle ones are kept, and how long an idle one may sit:
import httpx
limits = httpx.Limits(
max_connections=50, # total, across all hosts (default 100)
max_keepalive_connections=20, # idle connections kept for reuse (default 20)
keepalive_expiry=30.0, # seconds an idle connection is kept (default 5)
)
client = httpx.AsyncClient(limits=limits, timeout=httpx.Timeout(10.0, connect=3.0, pool=2.0))
max_connections is the concurrency cap on outbound requests from this client — every request beyond it waits for the pool. max_keepalive_connections decides how many connections survive a burst: tested, after a burst of 10 concurrent requests on a client with max_keepalive_connections=2, the pool held 2 connections, so the next burst paid 8 new connection setups. keepalive_expiry should be shorter than the server's idle timeout, or the client will try to reuse connections the server has already closed — the stale-connection problem in handling stale pooled connections after idle timeouts. Note that max_connections is a single total; httpx has no per-host limit, so one slow host can occupy the whole pool.
Verify: under a steady load, the number of open connections (ss -tn dst <host>) settles near your concurrency and below max_connections.
2. Understand what the pool timeout bounds¶
The pool timeout is how long a request waits for a connection when none becomes available. It did not cap the total time spent queueing while the pool kept turning over:
async with httpx.AsyncClient(limits=httpx.Limits(max_connections=5),
timeout=httpx.Timeout(5.0, pool=0.5)) as client:
await asyncio.gather(*(client.get(url) for _ in range(50)))
# server answers in 0.2 s: 50 ok, slowest waited 2.04 s, no PoolTimeout
# server answers in 0.6 s: 5 ok, 45 PoolTimeout at 0.5 s
With a 0.2 s server, a connection was released every 0.2 s — always inside the 0.5 s window — so waiting requests never timed out, even the one that queued behind nine rounds. With a 0.6 s server, nothing was released for 0.6 s, so every waiting request failed at 0.5 s. The pool timeout is therefore a detector for a stuck pool, not a bound on latency. For a latency bound, see the next step.
Verify: reproduce with your settings: a burst well above max_connections against a fast endpoint shows queueing far beyond pool without errors.
3. Bound total request time with asyncio.timeout¶
httpx's timeouts are all per phase — connect, each read, each write, pool wait. None bounds the whole request. Put a deadline around it:
async def fetch(client: httpx.AsyncClient, url: str, budget: float = 1.0) -> httpx.Response:
async with asyncio.timeout(budget): # includes pool wait, connect, transfer
return await client.get(url)
Tested with the same 5-connection pool and 0.2 s server: with a 1.0 s overall budget, 20 of 50 requests completed and 30 raised TimeoutError, instead of all 50 succeeding with up to 2 s of latency. Which is better depends on the caller: a user-facing request that has a deadline wants the fast failure; a batch job wants every request to finish. Make the choice explicit rather than inheriting it from the pool's behaviour.
Verify: under overload, the latency of successful requests never exceeds the budget.
4. Match max_connections to what the client is for¶
The right max_connections depends on the client's role. A client shared across a web service is a concurrency limiter for that dependency; a client in a batch job is a throughput knob:
# Web service: protect the dependency and fail fast when it is saturated
payments = httpx.AsyncClient(
base_url="https://payments.internal",
limits=httpx.Limits(max_connections=20, max_keepalive_connections=20, keepalive_expiry=30),
timeout=httpx.Timeout(3.0, pool=0.2), # a stuck pool fails within 0.2 s
)
# Batch job: maximize throughput up to what the target tolerates
bulk = httpx.AsyncClient(
limits=httpx.Limits(max_connections=50, max_keepalive_connections=50),
timeout=httpx.Timeout(30.0, pool=None), # wait as long as needed for a slot
)
Using one client per dependency gives each its own pool, so a slow dependency cannot exhaust connections that another one needs — a per-dependency bulkhead, as in bulkhead isolation with per-dependency semaphores. Set max_keepalive_connections equal to max_connections for steady high-concurrency traffic, so connections are not closed and reopened between bursts. For high in-flight counts, also check client CPU cost, as in choosing between httpx and aiohttp.
Verify: a slow dependency raises pool timeouts on its own client while requests to other dependencies continue at normal latency.
5. Watch the pool in production¶
httpx does not export pool metrics, but its trace extension reports each step of a request as it happens. The time from starting a request to sending its headers is the time spent waiting for and setting up a connection:
import time
async def get_with_pool_wait(client: httpx.AsyncClient, url: str) -> httpx.Response:
start = time.perf_counter()
waited = None
async def trace(event: str, info: dict) -> None:
nonlocal waited
if event == "http11.send_request_headers.started" and waited is None:
waited = time.perf_counter() - start
response = await client.get(url, extensions={"trace": trace})
POOL_WAIT.observe(waited or 0.0)
return response
Tested with six concurrent requests through a 2-connection pool to a 0.2 s endpoint, the measured waits were 0.00, 0.02, 0.21, 0.21, 0.41 and 0.41 s — the queue, made visible. (For HTTP/2 connections the event is named http2.send_request_headers.started.) Do not read response.elapsed in a response event hook: it raised RuntimeError there, because it is only available after the body has been read. A growing pool-wait figure is the earliest sign that max_connections is too low for the load or that the dependency has slowed down.
Verify: a load test that exceeds max_connections shows pool wait rising in the metric before errors appear.
Verification¶
httpx pools are configured well when:
- Limits are set per client role, with keepalive sized for steady traffic.
keepalive_expiryis below the server's idle timeout.- Total request time is bounded by
asyncio.timeoutwhere latency matters. - Pool wait is measured and alerts before timeouts occur.
Diagnostic Hook: compare the p99 of pool wait, measured with the trace extension, with the p99 of total request time. If pool wait grows with load while the time after the headers are sent stays flat, the client's pool is the bottleneck, not the server.
Pitfalls & edge cases¶
- Treating the pool timeout as a latency bound. Requests queued 2.04 s under a 0.5 s pool timeout.
- One client for every dependency. A slow one can take the whole pool.
- keepalive_expiry above the server's idle timeout. Reuse of closed connections.
- max_keepalive_connections far below max_connections. Connections churn between bursts.
Frequently Asked Questions¶
What does the httpx pool timeout do?
It bounds how long a request waits for a connection while none becomes available. In testing it did not limit total queueing when connections kept freeing up: requests waited up to 2.04 s under a 0.5 s pool timeout without error.
How do I set a total timeout for an httpx request?
httpx's timeouts are per phase. Wrap the call in async with asyncio.timeout(seconds) to bound the whole request, including time waiting for a pooled connection.
What are the default httpx connection limits?
max_connections=100, max_keepalive_connections=20 and keepalive_expiry=5 seconds, with a 5 second default for each timeout.
Does httpx have a per-host connection limit?
No. max_connections is a total for the client; use a separate client per dependency or a semaphore per host to isolate hosts.
Related¶
- Connection Pooling & Keep-Alive — up to the topic overview.
- Configuring aiohttp TCPConnector limits — the same decisions for aiohttp.
- Network I/O & Protocol Handling — the section overview.