Measuring TLS Handshake Cost in Async Services¶
TLS is cheap on an established connection and expensive to establish, and asyncio services pay the establishment cost in two places at once: the handshake's CPU work, which runs on the event loop thread, and the TCP accept queue, which overflows under bursts long before CPU does. Measured on Python 3.14 with OpenSSL 3.5.5, client and server on one machine, both on asyncio: a request on a new TLS connection took 1.30 ms with an RSA-2048 certificate and 1.17 ms with ECDSA P-256, almost all of it handshake CPU (median handshake 0.99 ms and 0.85 ms); the same request on a reused TLS connection took 0.03 ms; a new plain TCP connection took 0.17 ms. Separately, a burst of 200 simultaneous connects to asyncio.start_server with its default backlog=100 left 99 of them waiting about one second — for plain TCP and TLS alike — because the overflowing SYNs were dropped and retransmitted; with backlog=1024, the slowest of the 200 took 28 ms (TCP) and 253 ms (TLS). This guide measures both costs and keeps them off the critical path.
Prerequisites¶
- Python 3.11+,
pip install trustmefor test certificates. - asyncio servers, from writing a TCP server with asyncio.start_server.
- TLS on streams, from adding TLS to asyncio streams with SSL contexts.
1. Measure handshake cost against reuse¶
Generate a throwaway CA and certificate with trustme, serve TLS with start_server, and time a request on a new connection versus a reused one:
import asyncio, ssl, time
import trustme
ca = trustme.CA(key_type=trustme.KeyType.ECDSA)
cert = ca.issue_cert("localhost", key_type=trustme.KeyType.ECDSA)
server_ctx = ssl.create_default_context(ssl.Purpose.CLIENT_AUTH); cert.configure_cert(server_ctx)
client_ctx = ssl.create_default_context(); ca.configure_trust(client_ctx)
async def new_connection_request(port: int) -> float:
t = time.perf_counter()
reader, writer = await asyncio.open_connection("127.0.0.1", port, ssl=client_ctx,
server_hostname="localhost")
writer.write(b"ping\n"); await reader.readline()
writer.close(); await writer.wait_closed()
return time.perf_counter() - t
Measured over 300 requests each: 1.30 ms (RSA) and 1.17 ms (ECDSA) per request on a new TLS connection, of which the handshake itself was 0.99 ms and 0.85 ms at the median; 0.029–0.030 ms per request on one reused TLS connection; 0.17 ms on a new plain TCP connection. The process's CPU time per new-connection request matched its wall time — the handshake is computation, done on the event loop thread for both client and server in this test. ECDSA certificates were about 10–15% cheaper than RSA-2048. In production, add the network round trips the handshake needs: TLS 1.3 adds one round trip to TCP's, so on a 20 ms path a new connection costs about 40 ms before the first byte of the request.
Verify: your client's connection reuse rate is high — most requests should go over an existing connection.
2. Reuse connections instead of handshaking¶
The cheapest handshake is the one you do not do. Every async HTTP client pools connections, but only if the client object is reused:
# Each call handshakes: a new client, a new pool, a new TLS session
async def fetch_bad(url: str) -> bytes:
async with httpx.AsyncClient() as client:
return (await client.get(url)).content
# One client for the process: requests reuse pooled TLS connections
CLIENT = httpx.AsyncClient(timeout=httpx.Timeout(10.0, connect=3.0),
limits=httpx.Limits(max_keepalive_connections=50, keepalive_expiry=30))
async def fetch(url: str) -> bytes:
return (await CLIENT.get(url)).content
The per-call version pays the measured ~1.2 ms of handshake CPU plus the network round trips on every request, and creating the client itself cost a further 3 ms in testing because it builds a new SSL context, as measured in reusing SSL contexts in async clients. Keep-alive expiry should be shorter than the server's idle timeout so the client does not reuse connections the server has already closed — the stale-connection problem in handling stale pooled connections after idle timeouts.
Verify: connection count to each upstream stays roughly constant under steady load, rather than tracking request count.
3. Raise the accept backlog for bursty traffic¶
The second cost appears on the server, before TLS starts. asyncio.start_server and loop.create_server listen with backlog=100 by default. When more connections arrive at once than the queue holds, the kernel drops the excess SYNs and the clients retransmit after a second:
server = await asyncio.start_server(handler, host, port, ssl=server_ctx,
backlog=1024) # default is 100
Measured with 200 simultaneous connects: with the default backlog, 99 of 200 took more than 500 ms — about one second each, the SYN retransmission timeout — for both plain TCP (max 1,024 ms) and TLS (max 1,150 ms). With backlog=1024, the slowest plain connect took 28 ms and the slowest TLS connect 253 ms, the latter being handshake CPU queuing on one event loop. The kernel caps the backlog at net.core.somaxconn (4,096 on the test machine). Bursts of this size are ordinary: a fleet of clients reconnecting after a deploy, a load balancer opening connections to a new instance, a cron job's fan-out.
Verify: ss -ltn shows the listening socket's backlog (the Send-Q column for listeners), and a 200-connection burst completes without one-second outliers.
4. Keep handshake CPU from starving the loop¶
With a large backlog, a burst of TLS connections becomes a burst of handshake CPU on the event loop: 200 handshakes at roughly 1 ms each is 200 ms during which other requests wait. Spread the load and bound it:
# Client side: a pool limit bounds how many handshakes a burst of requests can start
client = httpx.AsyncClient(limits=httpx.Limits(max_connections=32, max_keepalive_connections=32))
# Server side: spread handshakes across cores with several processes
# uvicorn app:app --workers 4 --ssl-certfile cert.pem --ssl-keyfile key.pem
# or terminate TLS at the load balancer and serve plain HTTP to the Python process
For servers, the effective measures are more processes — handshakes then spread across cores, as in sizing uvicorn workers for async services — and terminating TLS at a load balancer or proxy that is built for it, so the Python process receives plain connections over a trusted network. For clients, the measure is a cap on concurrent new connections: a pool's max_connections bounds how many handshakes a burst of requests can trigger at once. An asyncio.Semaphore in the handler cannot help on the server side, because asyncio's transport performs the handshake before the handler runs.
Verify: during a connection burst, event-loop lag on the server stays within budget, or TLS terminates upstream of the Python process.
5. Monitor handshakes as a first-class metric¶
Handshake rate and duration tell you whether connection reuse is working, and where the time goes when it is not:
import time
async def timed_open(host: str, port: int, ctx: ssl.SSLContext):
t0 = time.perf_counter()
reader, writer = await asyncio.open_connection(host, port) # TCP only
t1 = time.perf_counter()
await writer.start_tls(ctx, server_hostname=host) # TLS only
t2 = time.perf_counter()
TCP_CONNECT.observe(t1 - t0)
TLS_HANDSHAKE.observe(t2 - t1)
return reader, writer
Separating TCP connect time from handshake time distinguishes a slow network or a full accept queue (TCP time, often in one-second steps) from CPU pressure or a slow certificate chain (handshake time). On the server side, the number of new TLS connections per second compared with requests per second is the reuse ratio; a ratio near one means clients are not keeping connections alive. In HTTP clients, aiohttp's trace hooks and httpx's event hooks expose similar timings, as in instrumenting httpx and aiohttp with OpenTelemetry.
Verify: dashboards show TCP connect time and TLS handshake time as separate series, plus new connections per request.
Verification¶
Connection setup is under control when:
- Clients reuse pooled connections, and new connections per request are near zero in steady state.
- Servers listen with a backlog sized for bursts, and bursts produce no one-second outliers.
- Handshake CPU is spread across processes or terminated upstream.
- TCP connect and TLS handshake times are measured separately.
Diagnostic Hook: look for one-second steps in connect-time histograms. A cluster of connects at about 1.0 s (and 3 s, 7 s for repeated losses) is SYN retransmission from an overflowing accept queue — a backlog problem, not a TLS one, and invisible in server-side request metrics because the requests never reached the application.
Pitfalls & edge cases¶
- The default backlog of 100. Measured: 99 of 200 burst connections waited about one second.
- A client per request. Every request pays the handshake, plus about 3 ms to build an SSL context.
- Handshakes on one loop. A burst serializes on the event loop's CPU.
- Measuring only total latency. TCP, TLS and request time need separate series.
Frequently Asked Questions¶
How expensive is a TLS handshake in Python asyncio?
About 0.85–0.99 ms of CPU for the handshake on loopback in testing (ECDSA and RSA certificates), making a request on a new connection 1.17–1.30 ms against 0.03 ms on a reused one — before network round trips.
Why do some connections to my asyncio server take exactly one second?
The listen backlog overflowed and the kernel dropped SYNs, which clients retransmit after about a second. asyncio's default backlog is 100; with 200 simultaneous connects, 99 waited about a second. Pass backlog=1024 to start_server.
Are ECDSA certificates faster than RSA for TLS?
In testing, an ECDSA P-256 handshake took 0.85 ms against 0.99 ms for RSA-2048, about 10–15% less.
How do I avoid TLS handshake overhead in async clients?
Create one client per process and reuse it, so requests go over pooled keep-alive connections, and set keep-alive expiry below the server's idle timeout.
Related¶
- TLS & DNS — up to the topic overview.
- Reusing SSL contexts in async clients — the cost before the handshake.
- Network I/O & Protocol Handling — the section overview.