Skip to content

Configuring httpx Limits and Pool Timeouts

An httpx.AsyncClient keeps a pool of connections, sized by httpx.Limits, and waits for a free one under the pool timeout. The defaults — 100 connections, 20 kept alive, 5 s for every timeout — are reasonable for many services and wrong for some. The pool timeout in particular does not mean what it seems to. Tested with httpx 0.28 against a server answering in 0.2 s, 50 concurrent requests through a pool limited to 5 connections with pool=0.5: none raised PoolTimeout, though the median request waited 1.23 s and the slowest 2.04 s — the timeout never fired because a connection became free every 0.2 s. With a 0.6 s server, 45 of 50 failed with PoolTimeout at 0.5 s. To bound how long a request may queue in total, wrap it in asyncio.timeout. This guide sets each limit deliberately and explains what each timeout does and does not bound.

Prerequisites

1. Set the three limits

httpx.Limits has three knobs: how many connections may exist, how many idle ones are kept, and how long an idle one may sit:

import httpx

limits = httpx.Limits(
    max_connections=50,            # total, across all hosts (default 100)
    max_keepalive_connections=20,  # idle connections kept for reuse (default 20)
    keepalive_expiry=30.0,         # seconds an idle connection is kept (default 5)
)
client = httpx.AsyncClient(limits=limits, timeout=httpx.Timeout(10.0, connect=3.0, pool=2.0))

max_connections is the concurrency cap on outbound requests from this client — every request beyond it waits for the pool. max_keepalive_connections decides how many connections survive a burst: tested, after a burst of 10 concurrent requests on a client with max_keepalive_connections=2, the pool held 2 connections, so the next burst paid 8 new connection setups. keepalive_expiry should be shorter than the server's idle timeout, or the client will try to reuse connections the server has already closed — the stale-connection problem in handling stale pooled connections after idle timeouts. Note that max_connections is a single total; httpx has no per-host limit, so one slow host can occupy the whole pool.

Verify: under a steady load, the number of open connections (ss -tn dst <host>) settles near your concurrency and below max_connections.

Request latency through a 5-connection pool 4 horizontal bars comparing fastest request with the others. Request latency through a 5-connection pool fastest request 0.21 s median request 1.23 s slowest request 2.04 s pool timeout setting 0.5 s, never fired httpx 0.28.1, Limits(max_connections=5), Timeout(5.0, pool=0.5); server answers in 0.2 s. Queueing time far exceeded the pool timeout because connections kept freeing up.

2. Understand what the pool timeout bounds

The pool timeout is how long a request waits for a connection when none becomes available. It did not cap the total time spent queueing while the pool kept turning over:

async with httpx.AsyncClient(limits=httpx.Limits(max_connections=5),
                             timeout=httpx.Timeout(5.0, pool=0.5)) as client:
    await asyncio.gather(*(client.get(url) for _ in range(50)))
    # server answers in 0.2 s: 50 ok, slowest waited 2.04 s, no PoolTimeout
    # server answers in 0.6 s: 5 ok, 45 PoolTimeout at 0.5 s

With a 0.2 s server, a connection was released every 0.2 s — always inside the 0.5 s window — so waiting requests never timed out, even the one that queued behind nine rounds. With a 0.6 s server, nothing was released for 0.6 s, so every waiting request failed at 0.5 s. The pool timeout is therefore a detector for a stuck pool, not a bound on latency. For a latency bound, see the next step.

Verify: reproduce with your settings: a burst well above max_connections against a fast endpoint shows queueing far beyond pool without errors.

3. Bound total request time with asyncio.timeout

httpx's timeouts are all per phase — connect, each read, each write, pool wait. None bounds the whole request. Put a deadline around it:

async def fetch(client: httpx.AsyncClient, url: str, budget: float = 1.0) -> httpx.Response:
    async with asyncio.timeout(budget):           # includes pool wait, connect, transfer
        return await client.get(url)

Tested with the same 5-connection pool and 0.2 s server: with a 1.0 s overall budget, 20 of 50 requests completed and 30 raised TimeoutError, instead of all 50 succeeding with up to 2 s of latency. Which is better depends on the caller: a user-facing request that has a deadline wants the fast failure; a batch job wants every request to finish. Make the choice explicit rather than inheriting it from the pool's behaviour.

Verify: under overload, the latency of successful requests never exceeds the budget.

What each httpx timeout bounds A grid of 5 rows by 3 columns. What each httpx timeout bounds timeout bounds does not bound connect TCP + TLS handshake anything after read gap between received chunks total body time write each send operation total upload time pool wait while no connection frees up total time queued asyncio.timeout(...) the whole call (it is the total) Per-phase timeouts catch stalls; only a deadline bounds latency.

4. Match max_connections to what the client is for

The right max_connections depends on the client's role. A client shared across a web service is a concurrency limiter for that dependency; a client in a batch job is a throughput knob:

# Web service: protect the dependency and fail fast when it is saturated
payments = httpx.AsyncClient(
    base_url="https://payments.internal",
    limits=httpx.Limits(max_connections=20, max_keepalive_connections=20, keepalive_expiry=30),
    timeout=httpx.Timeout(3.0, pool=0.2),     # a stuck pool fails within 0.2 s
)

# Batch job: maximize throughput up to what the target tolerates
bulk = httpx.AsyncClient(
    limits=httpx.Limits(max_connections=50, max_keepalive_connections=50),
    timeout=httpx.Timeout(30.0, pool=None),   # wait as long as needed for a slot
)

Using one client per dependency gives each its own pool, so a slow dependency cannot exhaust connections that another one needs — a per-dependency bulkhead, as in bulkhead isolation with per-dependency semaphores. Set max_keepalive_connections equal to max_connections for steady high-concurrency traffic, so connections are not closed and reopened between bursts. For high in-flight counts, also check client CPU cost, as in choosing between httpx and aiohttp.

Verify: a slow dependency raises pool timeouts on its own client while requests to other dependencies continue at normal latency.

5. Watch the pool in production

httpx does not export pool metrics, but its trace extension reports each step of a request as it happens. The time from starting a request to sending its headers is the time spent waiting for and setting up a connection:

import time


async def get_with_pool_wait(client: httpx.AsyncClient, url: str) -> httpx.Response:
    start = time.perf_counter()
    waited = None

    async def trace(event: str, info: dict) -> None:
        nonlocal waited
        if event == "http11.send_request_headers.started" and waited is None:
            waited = time.perf_counter() - start

    response = await client.get(url, extensions={"trace": trace})
    POOL_WAIT.observe(waited or 0.0)
    return response

Tested with six concurrent requests through a 2-connection pool to a 0.2 s endpoint, the measured waits were 0.00, 0.02, 0.21, 0.21, 0.41 and 0.41 s — the queue, made visible. (For HTTP/2 connections the event is named http2.send_request_headers.started.) Do not read response.elapsed in a response event hook: it raised RuntimeError there, because it is only available after the body has been read. A growing pool-wait figure is the earliest sign that max_connections is too low for the load or that the dependency has slowed down.

Verify: a load test that exceeds max_connections shows pool wait rising in the metric before errors appear.

Which pool settings fit this client? A decision on What is the client for with 3 outcomes. Which pool settings fit this client? What is the client for? dependency of a web service cap + short pool timeout plus asyncio.timeout batch or crawler cap by measured throughput pool=None several dependencies one client each bulkhead Limits express intent; set them per role, not once globally.

Verification

httpx pools are configured well when:

  • Limits are set per client role, with keepalive sized for steady traffic.
  • keepalive_expiry is below the server's idle timeout.
  • Total request time is bounded by asyncio.timeout where latency matters.
  • Pool wait is measured and alerts before timeouts occur.

Diagnostic Hook: compare the p99 of pool wait, measured with the trace extension, with the p99 of total request time. If pool wait grows with load while the time after the headers are sent stays flat, the client's pool is the bottleneck, not the server.

Pitfalls & edge cases

  • Treating the pool timeout as a latency bound. Requests queued 2.04 s under a 0.5 s pool timeout.
  • One client for every dependency. A slow one can take the whole pool.
  • keepalive_expiry above the server's idle timeout. Reuse of closed connections.
  • max_keepalive_connections far below max_connections. Connections churn between bursts.

Frequently Asked Questions

What does the httpx pool timeout do?

It bounds how long a request waits for a connection while none becomes available. In testing it did not limit total queueing when connections kept freeing up: requests waited up to 2.04 s under a 0.5 s pool timeout without error.

How do I set a total timeout for an httpx request?

httpx's timeouts are per phase. Wrap the call in async with asyncio.timeout(seconds) to bound the whole request, including time waiting for a pooled connection.

What are the default httpx connection limits?

max_connections=100, max_keepalive_connections=20 and keepalive_expiry=5 seconds, with a 5 second default for each timeout.

Does httpx have a per-host connection limit?

No. max_connections is a total for the client; use a separate client per dependency or a semaphore per host to isolate hosts.