Skip to content

Measuring TLS Handshake Cost in Async Services

TLS is cheap on an established connection and expensive to establish, and asyncio services pay the establishment cost in two places at once: the handshake's CPU work, which runs on the event loop thread, and the TCP accept queue, which overflows under bursts long before CPU does. Measured on Python 3.14 with OpenSSL 3.5.5, client and server on one machine, both on asyncio: a request on a new TLS connection took 1.30 ms with an RSA-2048 certificate and 1.17 ms with ECDSA P-256, almost all of it handshake CPU (median handshake 0.99 ms and 0.85 ms); the same request on a reused TLS connection took 0.03 ms; a new plain TCP connection took 0.17 ms. Separately, a burst of 200 simultaneous connects to asyncio.start_server with its default backlog=100 left 99 of them waiting about one second — for plain TCP and TLS alike — because the overflowing SYNs were dropped and retransmitted; with backlog=1024, the slowest of the 200 took 28 ms (TCP) and 253 ms (TLS). This guide measures both costs and keeps them off the critical path.

Prerequisites

1. Measure handshake cost against reuse

Generate a throwaway CA and certificate with trustme, serve TLS with start_server, and time a request on a new connection versus a reused one:

import asyncio, ssl, time
import trustme

ca = trustme.CA(key_type=trustme.KeyType.ECDSA)
cert = ca.issue_cert("localhost", key_type=trustme.KeyType.ECDSA)
server_ctx = ssl.create_default_context(ssl.Purpose.CLIENT_AUTH); cert.configure_cert(server_ctx)
client_ctx = ssl.create_default_context(); ca.configure_trust(client_ctx)


async def new_connection_request(port: int) -> float:
    t = time.perf_counter()
    reader, writer = await asyncio.open_connection("127.0.0.1", port, ssl=client_ctx,
                                                   server_hostname="localhost")
    writer.write(b"ping\n"); await reader.readline()
    writer.close(); await writer.wait_closed()
    return time.perf_counter() - t

Measured over 300 requests each: 1.30 ms (RSA) and 1.17 ms (ECDSA) per request on a new TLS connection, of which the handshake itself was 0.99 ms and 0.85 ms at the median; 0.029–0.030 ms per request on one reused TLS connection; 0.17 ms on a new plain TCP connection. The process's CPU time per new-connection request matched its wall time — the handshake is computation, done on the event loop thread for both client and server in this test. ECDSA certificates were about 10–15% cheaper than RSA-2048. In production, add the network round trips the handshake needs: TLS 1.3 adds one round trip to TCP's, so on a 20 ms path a new connection costs about 40 ms before the first byte of the request.

Verify: your client's connection reuse rate is high — most requests should go over an existing connection.

Cost per request by connection strategy 4 horizontal bars comparing new TLS conn, RSA-2048 with the others. Cost per request by connection strategy new TLS conn, RSA-2048 1.30 ms new TLS conn, ECDSA P-256 1.17 ms new plain TCP conn 0.17 ms reused TLS conn 0.03 ms Python 3.14, OpenSSL 3.5.5, client and server on one host; network round trips not included. The handshake is about forty times the cost of a request on a warm connection.

2. Reuse connections instead of handshaking

The cheapest handshake is the one you do not do. Every async HTTP client pools connections, but only if the client object is reused:

# Each call handshakes: a new client, a new pool, a new TLS session
async def fetch_bad(url: str) -> bytes:
    async with httpx.AsyncClient() as client:
        return (await client.get(url)).content


# One client for the process: requests reuse pooled TLS connections
CLIENT = httpx.AsyncClient(timeout=httpx.Timeout(10.0, connect=3.0),
                           limits=httpx.Limits(max_keepalive_connections=50, keepalive_expiry=30))

async def fetch(url: str) -> bytes:
    return (await CLIENT.get(url)).content

The per-call version pays the measured ~1.2 ms of handshake CPU plus the network round trips on every request, and creating the client itself cost a further 3 ms in testing because it builds a new SSL context, as measured in reusing SSL contexts in async clients. Keep-alive expiry should be shorter than the server's idle timeout so the client does not reuse connections the server has already closed — the stale-connection problem in handling stale pooled connections after idle timeouts.

Verify: connection count to each upstream stays roughly constant under steady load, rather than tracking request count.

3. Raise the accept backlog for bursty traffic

The second cost appears on the server, before TLS starts. asyncio.start_server and loop.create_server listen with backlog=100 by default. When more connections arrive at once than the queue holds, the kernel drops the excess SYNs and the clients retransmit after a second:

server = await asyncio.start_server(handler, host, port, ssl=server_ctx,
                                    backlog=1024)        # default is 100

Measured with 200 simultaneous connects: with the default backlog, 99 of 200 took more than 500 ms — about one second each, the SYN retransmission timeout — for both plain TCP (max 1,024 ms) and TLS (max 1,150 ms). With backlog=1024, the slowest plain connect took 28 ms and the slowest TLS connect 253 ms, the latter being handshake CPU queuing on one event loop. The kernel caps the backlog at net.core.somaxconn (4,096 on the test machine). Bursts of this size are ordinary: a fleet of clients reconnecting after a deploy, a load balancer opening connections to a new instance, a cron job's fan-out.

Verify: ss -ltn shows the listening socket's backlog (the Send-Q column for listeners), and a 200-connection burst completes without one-second outliers.

200 simultaneous connects to asyncio.start_server A grid of 4 rows by 5 columns. 200 simultaneous connects to asyncio.start_server server backlog over 500 ms p50 max plain TCP 100 (default) 99 of 200 24.8 ms 1,024 ms plain TCP 1024 0 25.2 ms 28.2 ms TLS 100 (default) 99 of 200 140.5 ms 1,150 ms TLS 1024 0 248.8 ms 252.6 ms Dropped SYNs are retransmitted after about a second; the backlog decides how many are dropped.

4. Keep handshake CPU from starving the loop

With a large backlog, a burst of TLS connections becomes a burst of handshake CPU on the event loop: 200 handshakes at roughly 1 ms each is 200 ms during which other requests wait. Spread the load and bound it:

# Client side: a pool limit bounds how many handshakes a burst of requests can start
client = httpx.AsyncClient(limits=httpx.Limits(max_connections=32, max_keepalive_connections=32))

# Server side: spread handshakes across cores with several processes
#   uvicorn app:app --workers 4 --ssl-certfile cert.pem --ssl-keyfile key.pem
# or terminate TLS at the load balancer and serve plain HTTP to the Python process

For servers, the effective measures are more processes — handshakes then spread across cores, as in sizing uvicorn workers for async services — and terminating TLS at a load balancer or proxy that is built for it, so the Python process receives plain connections over a trusted network. For clients, the measure is a cap on concurrent new connections: a pool's max_connections bounds how many handshakes a burst of requests can trigger at once. An asyncio.Semaphore in the handler cannot help on the server side, because asyncio's transport performs the handshake before the handler runs.

Verify: during a connection burst, event-loop lag on the server stays within budget, or TLS terminates upstream of the Python process.

5. Monitor handshakes as a first-class metric

Handshake rate and duration tell you whether connection reuse is working, and where the time goes when it is not:

import time


async def timed_open(host: str, port: int, ctx: ssl.SSLContext):
    t0 = time.perf_counter()
    reader, writer = await asyncio.open_connection(host, port)               # TCP only
    t1 = time.perf_counter()
    await writer.start_tls(ctx, server_hostname=host)                        # TLS only
    t2 = time.perf_counter()
    TCP_CONNECT.observe(t1 - t0)
    TLS_HANDSHAKE.observe(t2 - t1)
    return reader, writer

Separating TCP connect time from handshake time distinguishes a slow network or a full accept queue (TCP time, often in one-second steps) from CPU pressure or a slow certificate chain (handshake time). On the server side, the number of new TLS connections per second compared with requests per second is the reuse ratio; a ratio near one means clients are not keeping connections alive. In HTTP clients, aiohttp's trace hooks and httpx's event hooks expose similar timings, as in instrumenting httpx and aiohttp with OpenTelemetry.

Verify: dashboards show TCP connect time and TLS handshake time as separate series, plus new connections per request.

Where is connection setup time going? A decision on What does the measurement show with 4 outcomes. Where is connection setup time going? What does the measurement show? new connection per request reuse one pooled client 0.03 vs 1.2 ms 1 s outliers during bursts raise backlog (default 100) 99 of 200 waited loop lag during handshake bursts more processes or TLS at the proxy CPU on the loop handshake CPU dominates ECDSA certificates 10-15% cheaper Reuse removes handshakes; the backlog and processes absorb the ones that remain.

Verification

Connection setup is under control when:

  • Clients reuse pooled connections, and new connections per request are near zero in steady state.
  • Servers listen with a backlog sized for bursts, and bursts produce no one-second outliers.
  • Handshake CPU is spread across processes or terminated upstream.
  • TCP connect and TLS handshake times are measured separately.

Diagnostic Hook: look for one-second steps in connect-time histograms. A cluster of connects at about 1.0 s (and 3 s, 7 s for repeated losses) is SYN retransmission from an overflowing accept queue — a backlog problem, not a TLS one, and invisible in server-side request metrics because the requests never reached the application.

Pitfalls & edge cases

  • The default backlog of 100. Measured: 99 of 200 burst connections waited about one second.
  • A client per request. Every request pays the handshake, plus about 3 ms to build an SSL context.
  • Handshakes on one loop. A burst serializes on the event loop's CPU.
  • Measuring only total latency. TCP, TLS and request time need separate series.

Frequently Asked Questions

How expensive is a TLS handshake in Python asyncio?

About 0.85–0.99 ms of CPU for the handshake on loopback in testing (ECDSA and RSA certificates), making a request on a new connection 1.17–1.30 ms against 0.03 ms on a reused one — before network round trips.

Why do some connections to my asyncio server take exactly one second?

The listen backlog overflowed and the kernel dropped SYNs, which clients retransmit after about a second. asyncio's default backlog is 100; with 200 simultaneous connects, 99 waited about a second. Pass backlog=1024 to start_server.

Are ECDSA certificates faster than RSA for TLS?

In testing, an ECDSA P-256 handshake took 0.85 ms against 0.99 ms for RSA-2048, about 10–15% less.

How do I avoid TLS handshake overhead in async clients?

Create one client per process and reuse it, so requests go over pooled keep-alive connections, and set keep-alive expiry below the server's idle timeout.