TLS & DNS in asyncio¶
Every connection an asyncio service makes or accepts passes through two steps that are easy to treat as solved: resolving a name to addresses, and negotiating TLS. Both have asyncio-specific failure modes that show up as unexplained latency rather than errors. This section measures them on Python 3.14 with OpenSSL 3.5.5. loop.getaddrinfo runs in the default thread pool, so with that pool's 28 workers busy, lookups that normally take 1.4 ms took 951 ms; aiodns stayed at 1.7 ms. A first address that silently drops packets left a default asyncio.open_connection hanging past 10 s, while Happy Eyeballs connected in 0.25 s. A request on a new TLS connection cost 1.17–1.30 ms of CPU against 0.03 ms on a reused one, and a burst of 200 connects to a server with asyncio's default backlog of 100 left 99 of them waiting a second for SYN retransmission. Building an SSL context cost 2.7 ms, which made httpx.AsyncClient() take 3.06 ms to construct. Reloading a certificate into a live context took 0.09 ms — and a bad bundle broke every new handshake. And plaintext pipelined after STARTTLS reached the TLS layer and failed the handshake, unless the server checked for it first.
The parent section, Network I/O & Protocol Handling, covers streams, pools and HTTP clients; this topic covers what happens between "I have a hostname" and "I have an authenticated, encrypted connection". Each guide isolates one step, measures it, and gives the configuration that removes its failure mode.
Architectural principles¶
- Name resolution must not share a saturated thread pool. Use aiodns, or keep blocking work out of the default executor.
- Every connect races addresses and has a deadline. Happy Eyeballs for dead addresses, a timeout for when all are dead.
- Handshakes are expensive; connections are reused. One client per process, keep-alive, and a server backlog sized for bursts.
- SSL contexts are built once per trust decision and passed to every client explicitly.
- TLS boundaries are enforced on purpose: certificates swap only after validation, and nothing sent before a handshake is trusted after it.
Execution model: where each step runs¶
Each step runs somewhere different, and that is where its costs come from. Resolution runs socket.getaddrinfo in the loop's default ThreadPoolExecutor — min(32, cpu_count + 4) threads, shared with every asyncio.to_thread and run_in_executor(None, ...) call — unless a client uses an asynchronous resolver such as aiodns, which registers c-ares sockets with the event loop instead. httpx resolves through anyio, which calls loop.getaddrinfo too; aiohttp's default ThreadedResolver does the same, its AsyncResolver does not.
Connecting is non-blocking I/O on the loop, but the kernel decides how long an unanswered SYN waits: with Linux's default tcp_syn_retries = 6, about 127 seconds. Happy Eyeballs starts the next address after a delay; a timeout bounds the whole attempt. On the server, completed connections wait in the listen socket's accept queue, whose length is the backlog argument — 100 by default in asyncio — capped by net.core.somaxconn. TLS is CPU work on the event loop thread: the handshake runs inside the transport, before the protocol or stream handler sees the connection, so a burst of handshakes is a burst of loop time. Contexts are consulted per handshake, which is what makes live certificate swaps possible.
Pattern catalogue¶
Resolution off the shared pool¶
resolver = aiodns.DNSResolver()
result = await resolver.getaddrinfo(host, port=443, type=socket.SOCK_STREAM)
connector = aiohttp.TCPConnector(resolver=AsyncResolver(), ttl_dns_cache=300)
With the default executor saturated by 28 one-second jobs, thread-based lookups waited 951 ms while aiodns answered in 1.69 ms; with the pool idle the two were within half a millisecond of each other. The alternative to changing the resolver is moving heavy blocking work — file I/O, synchronous SDKs — onto dedicated executors so the default pool stays free for short tasks such as lookups. aiodns does not consult NSS modules, so hosts resolved through LDAP or mDNS need the system resolver. See resolving DNS without blocking the executor.
Racing addresses with a deadline¶
async with asyncio.timeout(3.0):
reader, writer = await asyncio.open_connection(host, port, happy_eyeballs_delay=0.25, interleave=1)
aiohttp and httpx race by default; raw asyncio connections do not. Measured with a black-holed first address: a default open_connection was still waiting at the 10 s test limit, happy_eyeballs_delay=0.25 connected in 0.25 s, aiohttp's default connector in 0.25 s and httpx in 0.33 s. The timeout matters independently, because when every address is dead the kernel's SYN retries take about two minutes. Protocol clients and database drivers built on raw asyncio connections are where the racing default is most often missing. See connecting with happy eyeballs in asyncio.
Reuse and backlog¶
One pooled client per process; servers listen with backlog=1024. A request on a reused TLS connection cost 0.03 ms against 1.17–1.30 ms on a new one, so connection reuse is the largest single saving in this section; the backlog fix is a one-argument change that removed every one-second outlier from a 200-connection burst. ECDSA certificates trimmed handshake CPU by 10–15% against RSA-2048. See measuring TLS handshake cost in async services.
Shared, named SSL contexts¶
PUBLIC = ssl.create_default_context(cafile=certifi.where())
http = httpx.AsyncClient(verify=PUBLIC) # 0.02 ms instead of 3.06 ms
A default context costs about 2.7 ms to build because it parses the whole trust store — 121 CA certificates on the test system — so the rule is one context per trust decision (public APIs, an internal CA, mutual TLS), named, built at startup and passed to every client and driver explicitly. Contexts are safe to share across tasks and clients.
See reusing SSL contexts in async clients.
Validated certificate swaps¶
Build and validate a new context, then switch handshakes to it through sni_callback. The swap is a single reference assignment, so every handshake sees either the old context or the new one, never a half-loaded state; the callback ran for clients with and without SNI. See reloading TLS certificates without restarting.
STARTTLS with an enforced boundary¶
Reject buffered bytes before start_tls; refuse plaintext fallback on the client. On Python 3.14 a pipelined command did not execute — the handshake failed instead — but an explicit check turns an accidental failure into a deliberate protocol error and keeps the guarantee independent of asyncio internals. See upgrading connections with STARTTLS in asyncio.
Choosing where TLS terminates¶
Many of the server-side concerns in this section — handshake CPU on the event loop, certificate rotation, accept-queue sizing — disappear from the Python process if TLS terminates at a load balancer or reverse proxy in front of it. That is often the right design: proxies are built for handshake throughput, rotate certificates through their own tooling, and present the application with plain, long-lived connections over a trusted network. Terminating TLS in the Python process is the right choice when there is no proxy (a standalone service, an edge agent), when the protocol is not HTTP and the proxy cannot speak it, or when mutual TLS identity must be checked by the application itself. In those cases the patterns here apply in full.
Client-side TLS has no such escape: every outbound call to an external API is the application's own handshake, and resolution, Happy Eyeballs, context reuse and connection reuse are the application's responsibility. The common production mix — TLS terminated at the proxy for inbound traffic, many outbound TLS clients from the application — puts most of this section's weight on the client side.
Testing connection setup deterministically¶
Most of these failure modes only appear under conditions that are hard to arrange by accident — a dead address, a full thread pool, a burst of connects, a malformed certificate bundle — so they deserve tests that arrange them on purpose. Three building blocks cover nearly everything in this section:
import trustme
CA = trustme.CA() # a throwaway certificate authority
SERVER_CERT = CA.issue_cert("localhost")
def server_context() -> ssl.SSLContext:
ctx = ssl.create_default_context(ssl.Purpose.CLIENT_AUTH)
SERVER_CERT.configure_cert(ctx)
return ctx
def client_context() -> ssl.SSLContext:
ctx = ssl.create_default_context()
CA.configure_trust(ctx)
return ctx
def dead_first_resolver(loop, host_name: str, live: str = "127.0.0.1"):
real = loop.getaddrinfo
async def fake(host, port, *args, **kwargs):
if host in (host_name, host_name.encode()):
return [(socket.AF_INET, socket.SOCK_STREAM, 6, "", ("10.255.255.1", port)),
(socket.AF_INET, socket.SOCK_STREAM, 6, "", (live, port))]
return await real(host, port, *args, **kwargs)
return fake
With a local CA, every TLS behaviour — handshake cost, certificate swaps, broken bundles, STARTTLS — can be exercised against a real asyncio.start_server in a unit test, without network access or real certificates. The fake resolver reproduces dead addresses (remember that some clients pass the hostname as bytes). A full thread pool is simply [asyncio.to_thread(time.sleep, 1.0) for _ in range(workers)] started before the lookups under test. Every measurement in this section came from scripts built on these pieces, and each one translates directly into a regression test with an assertion on time or outcome.
Resource boundaries¶
- Default executor threads:
min(32, cpu_count + 4), shared by DNS and everyto_thread. Keep heavy blocking work on dedicated executors. - Connect timeout: a few seconds, separate from read timeouts, because the kernel's own limit is about two minutes.
- Listen backlog: at least the largest expected burst of simultaneous connects, up to
somaxconn. - Handshake CPU: about 1 ms per handshake on the loop thread; bursts need processes or upstream termination.
- Handshake timeout: asyncio's default is 60 s per stalled peer; servers facing the internet want about 10.
- SSL contexts: one per trust decision, built at startup; their count should not grow with traffic.
Integrated production example¶
A TLS echo service and its client, combining the server-side and client-side patterns: validated certificate swaps on SIGHUP, a handshake timeout, a backlog sized for bursts, a shared client context, Happy Eyeballs and an overall deadline. In a test it served 300 concurrent TLS requests in 0.57 s, kept serving the old certificate when a renewal paired a new certificate with the wrong key, and presented the new certificate after a correct renewal:
import asyncio
import logging
import signal
import ssl
log = logging.getLogger("tls")
class CertStore:
"""Validated certificate swaps: handshakes always see a complete context."""
def __init__(self, cert: str, key: str) -> None:
self.cert, self.key = cert, key
self.current = self._build()
def _build(self) -> ssl.SSLContext:
ctx = ssl.create_default_context(ssl.Purpose.CLIENT_AUTH)
ctx.minimum_version = ssl.TLSVersion.TLSv1_2
ctx.load_cert_chain(self.cert, self.key)
return ctx
def reload(self) -> bool:
try:
self.current = self._build()
except (ssl.SSLError, OSError) as exc:
log.error("certificate reload rejected: %s", exc)
return False
log.info("certificate reloaded")
return True
def listening_context(self) -> ssl.SSLContext:
ctx = self._build()
ctx.sni_callback = lambda sslobj, name, _ctx: setattr(sslobj, "context", self.current)
return ctx
async def handle(reader: asyncio.StreamReader, writer: asyncio.StreamWriter) -> None:
try:
while line := await asyncio.wait_for(reader.readline(), timeout=30):
writer.write(line) # echo
await writer.drain()
except (TimeoutError, ConnectionError):
pass
finally:
writer.close()
async def serve(store: CertStore, host: str, port: int) -> asyncio.Server:
loop = asyncio.get_running_loop()
loop.add_signal_handler(signal.SIGHUP, store.reload)
return await asyncio.start_server(
handle, host, port,
ssl=store.listening_context(),
ssl_handshake_timeout=10.0, # stalled handshakes do not hold tasks for 60 s
backlog=1024, # bursts do not overflow into 1 s SYN retries
)
CLIENT_CTX = ssl.create_default_context() # built once: about 2.7 ms
async def request(host: str, port: int, payload: bytes, ctx: ssl.SSLContext = CLIENT_CTX) -> bytes:
async with asyncio.timeout(5.0): # connect + handshake + reply
reader, writer = await asyncio.open_connection(
host, port, ssl=ctx, server_hostname=host,
happy_eyeballs_delay=0.25, interleave=1)
try:
writer.write(payload + b"\n")
await writer.drain()
return (await reader.readline()).rstrip(b"\n")
finally:
writer.close()
The broken renewal in the test logged certificate reload rejected: [X509: KEY_VALUES_MISMATCH] and the next request still received its echo; the correct renewal logged certificate reloaded and the next connection presented the new certificate's serial. The client opens a connection per request for clarity; a real client that talks to the same server repeatedly should keep its connection or use a pooled HTTP client, saving the measured ~1.2 ms handshake per request.
Diagnostic hook callout¶
Measure each step separately, because the symptoms overlap:
- DNS time per lookup — spikes that track busy threads in the default executor mean lookups are queueing for threads.
- TCP connect time — clusters at about 1 s (and 3 s, 7 s) are SYN retransmissions from an overflowing accept queue; clusters at the Happy Eyeballs delay mean a dead first address is being skipped.
- TLS handshake time and rate — rising time with loop lag is CPU pressure; a handshake rate close to the request rate means connections are not being reused.
- Certificate expiry and reload failures — alert at seven days to expiry and on any rejected reload.
Alert on DNS p99 above 50 ms, any connect-time cluster near 1 s, and handshakes per request above about 0.1 for clients that should be pooling.
Failure modes¶
| Failure mode | Root cause | Detection | Fix |
|---|---|---|---|
| Connects slow only under load | DNS waiting for default-executor threads | DNS time tracks to_thread usage |
aiodns, or dedicated executors |
| Connections hang for minutes | Dead first address, no racing or timeout | Connect time near kernel limits | happy_eyeballs_delay, connect timeout |
| Some connects take exactly 1 s | Listen backlog overflow | 1 s clusters in connect time | backlog=1024 |
| High CPU per request on clients | New connection (and context) per request | Handshakes ≈ requests | One pooled client, shared context |
| Every new connection fails after renewal | Bad bundle loaded into the live context | Handshake errors after reload | Validate a new context, swap via sni_callback |
| Injected commands after STARTTLS | Buffered plaintext trusted after upgrade | Pipelining test | Reject buffered bytes before start_tls |
Frequently Asked Questions¶
Does asyncio resolve DNS asynchronously?
Not by itself: loop.getaddrinfo runs the blocking getaddrinfo in the default thread pool, so it waits when that pool is busy — 951 ms per lookup in testing with all 28 workers occupied. aiodns resolves without threads.
Why does my asyncio client hang when connecting?
Usually the first resolved address drops packets and the connect waits for the kernel's SYN retries. Pass happy_eyeballs_delay=0.25 to race addresses and wrap the connect in asyncio.timeout.
How expensive is TLS in asyncio?
About 1 ms of CPU per handshake on loopback, making a request on a new connection 1.17–1.30 ms against 0.03 ms on a reused one, plus network round trips. Reuse connections.
What backlog should an asyncio server use?
At least the largest burst of simultaneous connects; the default of 100 made 99 of 200 burst connections wait about a second. backlog=1024 removed the outliers.
How do I rotate certificates in an asyncio TLS server?
Build and validate a new SSLContext, then make new handshakes use it through an sni_callback on the listening context. Loading a bad bundle into the live context broke new connections in testing.
Related¶
- Resolving DNS without blocking the executor — lookups that never wait for a thread.
- Connecting with happy eyeballs in asyncio — around dead addresses.
- Measuring TLS handshake cost in async services — reuse and backlog.
- Reusing SSL contexts in async clients — build once, pass everywhere.
- Reloading TLS certificates without restarting — validated swaps.
- Upgrading connections with STARTTLS in asyncio — the plaintext boundary.
- Connection Pooling & Keep-Alive — keeping connections once they exist.
- Network I/O & Protocol Handling — the parent section.