Skip to content

Reusing Connections Across Lambda Invocations

Reusing connections between invocations is the main performance reason to keep async clients alive in a Lambda execution environment: a warm invocation that reuses a pooled connection skips the TCP connect and the TLS handshake. Between invocations, though, the environment is frozen — no code runs, no keep-alive pings are sent — while the servers on the other end keep their idle timers running. A pooled connection can therefore be closed by the server while the client is asleep, and the next invocation finds out when it tries to use it. Measured with the AWS Lambda Python 3.13 base image and its Runtime Interface Emulator, a handler with a module-level event loop and aiohttp session calling an nginx upstream whose keepalive_timeout was set to 2 s: the first invocation took 4.92 ms (new connection), the second 0.45 ms (reused). The container was then frozen with docker pause for 4 s — past the server's timeout — and resumed. The next invocation succeeded in 2.35 ms: aiohttp detected the closed connection and opened a new one without surfacing an error; the following invocations were back at 0.56–0.72 ms. This guide makes reuse safe when it works and harmless when it does not.

Prerequisites

1. Measure reuse across invocations

Return the handler's own timing so each invocation's cost is visible from outside:

LOOP = asyncio.new_event_loop()
SESSION: aiohttp.ClientSession | None = None


async def _fetch() -> int:
    global SESSION
    if SESSION is None:
        SESSION = aiohttp.ClientSession(connector=aiohttp.TCPConnector(keepalive_timeout=60))
    async with SESSION.get(UPSTREAM) as resp:
        await resp.read()
        return resp.status


def handler(event, context):
    t = time.perf_counter()
    status = LOOP.run_until_complete(_fetch())
    return {"status": status, "ms": round((time.perf_counter() - t) * 1e3, 2)}

Measured: 4.92 ms for the first invocation, which created the session and its first connection, then 0.45 ms for the second, which reused it. The upstream was a plain-HTTP nginx on the same Docker network, so the saving shown is only TCP setup and session creation; with TLS to a remote API, each avoided connection also saves a handshake and its round trips.

Verify: warm invocations are consistently faster than the first one in the same environment.

Handler time per invocation, with a 4 s freeze after invocation 2 5 horizontal bars comparing invoke 1: new session + connection with the others. Handler time per invocation, with a 4 s freeze after invocation 2 invoke 1: new session + connection 4.92 ms invoke 2: reused 0.45 ms invoke 3, after 4 s frozen: reconnected 2.35 ms invoke 4: reused 0.72 ms invoke 5: reused 0.56 ms AWS Lambda Python 3.13 image with the Runtime Interface Emulator; nginx keepalive_timeout 2 s; docker pause used to freeze. The stale connection cost one reconnect, not an error.

2. Expect connections to die while frozen

The freeze is invisible to the client: when the environment thaws, a pooled socket looks open until it is used. What happens next depends on the client and the request:

connector = aiohttp.TCPConnector(
    keepalive_timeout=30,          # client-side idle limit; shorter than most servers' and LBs'
    limit=10,
    enable_cleanup_closed=True,
)

Measured: after a 4-second freeze past nginx's 2-second keep-alive timeout, aiohttp's next GET completed normally in 2.35 ms — the closed connection was detected and replaced. That behaviour is client-specific, and it is safest for idempotent requests: a client that has already written a request on a dead connection cannot know whether the server processed it, so libraries retry such failures for safe methods but may raise for others. Setting the client's keep-alive timeout below the server's (and below any load balancer's idle timeout — often 60 seconds for AWS load balancers) makes the client drop idle connections before the server does, as long as the client's own clock runs; during a freeze it does not, so the first request after a long freeze is the one at risk.

Verify: a test that freezes the function longer than the upstream's idle timeout, then invokes it, succeeds for each kind of request the function makes.

3. Retry non-idempotent requests deliberately

For writes, make the retry explicit and safe rather than relying on the client:

async def post_order(session: aiohttp.ClientSession, order: dict) -> dict:
    key = order["idempotency_key"]                     # same key on every attempt
    for attempt in range(2):
        try:
            async with session.post(ORDERS, json=order,
                                    headers={"Idempotency-Key": key}) as resp:
                resp.raise_for_status()
                return await resp.json()
        except (aiohttp.ServerDisconnectedError, aiohttp.ClientOSError):
            if attempt == 1:
                raise
            # a pooled connection died while the environment was frozen; try once more
    raise AssertionError("unreachable")

An idempotency key lets the server deduplicate if the first attempt was in fact processed, which turns "retry after a dead connection" into a safe operation even for payments — the pattern in idempotency keys for safe async retries. One retry is enough for the stale-connection case; it is not a general retry policy.

Verify: a write retried after a forced connection drop is applied exactly once on the server.

The first request after a long freeze A sequence of 6 messages between 3 participants. The first request after a long freeze handler connection pool upstream server environment frozen: no keep-alive traffic idle timeout (2 s): FIN thaw, GET via pooled socket socket closed: discard, reconnect new connection + GET 200 (2.35 ms) Safe and transparent for GET; for writes, retry with an idempotency key.

4. Bound pool sizes to one invocation's needs

A Lambda environment handles one invocation at a time, so the connection pool never needs more connections than a single invocation uses concurrently:

connector = aiohttp.TCPConnector(
    limit=10,                  # concurrent requests within one invocation
    limit_per_host=10,
    keepalive_timeout=30,
)

Large pools in Lambda are pure overhead: each kept-alive connection holds a socket and buffers, and connections that are never reused just wait to be closed by the server. More importantly, every concurrent environment holds its own pool, so the upstream sees environments × pool size connections — a function scaled to 500 concurrent environments with a pool of 10 can hold 5,000 connections to a database. For databases specifically, that is why Lambda deployments usually put a proxy (RDS Proxy, PgBouncer) between functions and the database, as discussed in sizing async connection pools for throughput.

Verify: the upstream's connection count at peak concurrency stays within its limit, given the number of concurrent environments.

5. Close nothing at the end of an invocation

It is tempting to close the session at the end of each invocation to be tidy. That discards the pool — and the reason for keeping the session at all:

def handler(event, context):
    result = LOOP.run_until_complete(_work(event))
    # Do not: LOOP.run_until_complete(SESSION.close())  -> next invocation reconnects
    return result

Lambda gives no reliable hook at environment shutdown for releasing these resources; when an environment is retired, its sockets are torn down with it, and servers see ordinary connection closes. Leave pooled resources open between invocations, close per-invocation resources (files, temporary clients) within the invocation, and let the retirement clean up the rest. If a function does need cleanup at shutdown, Lambda extensions can receive a shutdown event, which is beyond the scope of a handler.

Verify: consecutive warm invocations do not create new sessions or pools, checked by a creation counter in logs.

How should this function treat pooled connections? A decision on What kind of call is it with 4 outcomes. How should this function treat pooled connections? What kind of call is it? idempotent GET client reconnects after a freeze 2.35 ms once write one retry with an idempotency key applied exactly once database from many environments small pools + a proxy environments x pool size any keep the session open between calls reuse is the point Reuse is an optimization that must fail safe.

Verification

Connection reuse across invocations is working when:

  • Warm invocations reuse pooled connections, visibly faster than cold ones.
  • The first request after a long freeze succeeds, and writes retry safely with idempotency keys.
  • Pool sizes match one invocation's concurrency, and upstream connection counts fit at peak scale.
  • Sessions are not closed between invocations.

Diagnostic Hook: log, per invocation, whether the session was new and how many new connections were opened (aiohttp's trace hooks expose on_connection_create_end). A new connection on most warm invocations means reuse is not happening — a session closed per invocation, or keep-alive timeouts shorter than the gap between invocations.

Pitfalls & edge cases

  • Assuming pooled connections survive freezes. Measured: the server closed one during a 4 s freeze.
  • Retrying writes blindly. Without an idempotency key, a retry can duplicate a write.
  • Large pools per environment. Multiplied by concurrent environments, they overwhelm databases.
  • Closing the session per invocation. It removes the reuse the module-level session exists for.

Frequently Asked Questions

Do HTTP connections survive between Lambda invocations?

They stay in the pool, but the server may close them while the environment is frozen. In testing, after a 4 s freeze past a 2 s server keep-alive timeout, aiohttp reconnected transparently on the next invocation (2.35 ms) and then reused the new connection.

Why does my Lambda get ServerDisconnectedError after being idle?

A pooled connection was closed by the server or a load balancer during the freeze. Keep client keep-alive below the server's idle timeout, retry idempotent requests once, and use idempotency keys for writes.

How big should a connection pool be in Lambda?

Big enough for one invocation's concurrent requests: an environment handles one invocation at a time, and each concurrent environment holds its own pool.

Should I close the aiohttp session at the end of a Lambda handler?

No, if you keep it at module level for reuse: closing it discards the pooled connections and the next invocation reconnects. Lambda tears the sockets down when it retires the environment.