Skip to content

Warming Connection Pools at Startup

A freshly started process has empty connection pools, so its first requests pay for connection setup: TCP, TLS, authentication, and driver initialization. Under a rolling deploy or a scale-up, those first requests arrive all at once — every new instance's first burst is slow, in the same seconds. Measured against PostgreSQL 17 on the same host, the first burst of 10 concurrent queries through an asyncpg pool created with min_size=0 had a median latency of 111 ms; the same burst through a pool created with min_size=10 took 1.4 ms, because the 108 ms of connection setup had already happened before traffic arrived. For httpx against a local HTTP server, the first burst's median was 9.7 ms cold and 3.2 ms warm. Across a network with TLS the cold cost is larger still. This guide opens connections during startup, ties readiness to it, and keeps warm-up from becoming a thundering herd of its own.

Prerequisites

1. Open database connections at pool creation

Most database pools have a minimum size that is opened when the pool is created. Set it to the connections you expect to need at steady load, and create the pool in startup:

@asynccontextmanager
async def lifespan(app):
    # min_size connections are opened here, before the app accepts traffic
    app.state.db = await asyncpg.create_pool(dsn, min_size=10, max_size=20)
    yield
    await app.state.db.close()

Measured: create_pool with min_size=10 took 108 ms — that is the cost moved out of the request path — and the first burst of 10 queries then ran at 1.4 ms median against 111 ms for the cold pool. The second burst was under 1 ms in both cases; warm-up only changes the first seconds, but those are the seconds when a new instance is most likely to be overwhelmed. psycopg's AsyncConnectionPool behaves the same way with min_size and await pool.open(wait=True), and SQLAlchemy's pool has no minimum, so warm it explicitly as in step 3.

Verify: immediately after startup, the database shows min_size connections from the new process before any traffic.

Median latency of the first burst of 10 queries 3 horizontal bars comparing cold pool (min_size=0) with the others. Median latency of the first burst of 10 queries cold pool (min_size=0) 111 ms warm pool (min_size=10) 1.4 ms any later burst 0.6 ms asyncpg 0.31, PostgreSQL 17 on localhost; the warm pool spent 108 ms connecting during startup. Warm-up does not remove the cost; it moves it before the first request.

2. Warm HTTP clients by touching each upstream

HTTP clients have no minimum pool size: connections open on first use. Warm them by making cheap requests to each upstream during startup, concurrently up to the connections you want kept:

async def warm_http(client: httpx.AsyncClient, url: str, n: int) -> None:
    async def touch():
        try:
            await client.head(url, timeout=2.0)
        except httpx.HTTPError:
            pass                           # a failed warm-up is not a failed startup

    await asyncio.gather(*(touch() for _ in range(n)))


@asynccontextmanager
async def lifespan(app):
    app.state.http = httpx.AsyncClient(limits=httpx.Limits(max_keepalive_connections=10))
    await warm_http(app.state.http, "https://payments.internal/healthz", n=10)
    yield
    await app.state.http.aclose()

Measured locally, the first burst of 10 requests had a 9.7 ms median and a 20.3 ms maximum cold, and 3.2 / 4.3 ms warm. Across regions with TLS, each new connection costs one or more extra round trips, so the gap is tens to hundreds of milliseconds. The warmed connections only survive until keepalive_expiry, so warm-up must happen shortly before traffic arrives — and max_keepalive_connections must be at least n, or the extra connections are closed as soon as they are idle.

Verify: right after startup, ss -tn shows n established connections to each warmed upstream.

3. Warm SQLAlchemy engines and caches explicitly

SQLAlchemy's async engine opens connections lazily and has no minimum size. Check out connections concurrently once and return them:

from sqlalchemy import text


async def warm_engine(engine, n: int) -> None:
    async def one():
        async with engine.connect() as conn:
            await conn.execute(text("select 1"))
            await asyncio.sleep(0.05)              # hold briefly so the checkouts overlap
    await asyncio.gather(*(one() for _ in range(n)))

Holding each connection briefly forces the pool to open n distinct connections rather than reusing one; once returned, they stay in the pool up to pool_size. The same idea applies to other lazy state: a first request that compiles templates, loads a model, or fills an in-process cache. Run those during startup too, as in warming caches on startup without blocking readiness.

Verify: the engine's pool.checkedin() reports n connections after warm-up.

Startup order for a warm instance A flow of 5 stages. Startup order for a warm instance create pools min_size opened touch upstreams HEAD / select 1 load lazy state caches, models mark ready probe returns 200 traffic arrives connections warm Readiness should mean "warm", not merely "process started".

4. Tie readiness to warm-up

Warm-up only helps if traffic waits for it. The orchestrator routes requests to an instance when its readiness probe passes, so the probe must not pass until warm-up is done:

class Readiness:
    def __init__(self) -> None:
        self.ready = asyncio.Event()


@asynccontextmanager
async def lifespan(app):
    app.state.readiness = Readiness()
    app.state.db = await asyncpg.create_pool(dsn, min_size=10, max_size=20)
    app.state.http = httpx.AsyncClient()
    await warm_http(app.state.http, UPSTREAM_HEALTH, n=10)
    app.state.readiness.ready.set()           # only now will the probe pass
    yield
    app.state.readiness.ready.clear()         # stop receiving traffic during shutdown
    await app.state.http.aclose()
    await app.state.db.close()


async def readyz(request):
    ok = request.app.state.readiness.ready.is_set()
    return Response(status_code=200 if ok else 503)

With Uvicorn, the lifespan startup completes before the server accepts connections at all, so anything done before yield already delays traffic; an explicit readiness flag matters when warm-up runs in the background or when a separate probe port is served earlier. Clearing the flag at shutdown takes the instance out of rotation before its pools close, the start of a graceful shutdown as in Graceful Shutdown & Signals.

Verify: during a rolling deploy, the p99 latency of the first minute on new instances matches steady state.

5. Keep warm-up from becoming a stampede

If every instance in a fleet starts at once — a full restart, a large scale-up — warm-up opens every connection at the same moment, and the database or upstream sees a connection storm. Bound and spread it:

import random


async def warm_with_jitter(pool_factory, max_delay: float = 2.0):
    await asyncio.sleep(random.uniform(0, max_delay))     # spread instances apart
    return await pool_factory()


@asynccontextmanager
async def lifespan(app):
    app.state.db = await warm_with_jitter(
        lambda: asyncpg.create_pool(dsn, min_size=5, max_size=20)   # warm a base, grow on demand
    )
    yield
    await app.state.db.close()

Warm a base of connections rather than the maximum, and let the pool grow under real load. A random delay of a second or two spreads a fleet's connection setup over time, and warm-up failures should never fail startup — log them and let the pool connect on demand, so a brief dependency outage cannot prevent the service from starting. Authentication and TLS handshakes are CPU-heavy on the server side, so this matters more for databases than for HTTP upstreams.

Verify: a simultaneous restart of every instance does not push the database's connection rate or CPU above its normal peak.

How should this pool be warmed? A decision on What kind of pool is it with 3 outcomes. How should this pool be warmed? What kind of pool is it? DB pool with min_size set min_size opens at creation lazy pool (SQLAlchemy, HTTP) concurrent touch requests at startup whole fleet restarts jitter + base size avoid a storm Warm enough to serve the first burst; let real load do the rest.

Verification

Warm-up is effective when:

  • Pools hold connections before the first request, at the expected steady-state size.
  • Readiness passes only after warm-up, and fails at shutdown.
  • First-minute latency on new instances matches steady state.
  • Fleet-wide restarts do not cause connection storms.

Diagnostic Hook: chart p99 latency per instance by age since start. A spike in the first seconds of each instance's life is cold pools or lazy state; compare it with the number of connections each instance holds at the moment it turned ready. If the instance is warm but still slow, look for other lazy initialization in the first request path.

Pitfalls & edge cases

  • Readiness that passes before warm-up. Traffic arrives on cold pools anyway.
  • Warming more connections than max_keepalive_connections. They are closed as soon as they idle.
  • Warming long before traffic. Connections expire before they are used.
  • Failing startup on warm-up errors. A dependency blip should not stop the service from starting.

Frequently Asked Questions

How do I pre-open database connections in an asyncio app?

Set the pool's minimum size and create it in the startup hook. In testing, an asyncpg pool with min_size=10 served its first burst at 1.4 ms median against 111 ms for a pool that connected on demand.

How do I warm up an httpx or aiohttp connection pool?

Make cheap concurrent requests, such as HEAD to a health endpoint, to each upstream during startup, with keepalive limits at least as large as the number of connections you warm.

Should readiness wait for pool warm-up?

Yes. Mark the instance ready only after pools and lazy state are warm, so the orchestrator does not send traffic to cold connections.

Can warming pools overload the database?

If a whole fleet starts at once, yes. Add random jitter before warming, warm a base size rather than the maximum, and never fail startup because warm-up failed.