Skip to content

Resolving DNS Without Blocking the Executor

Every outbound connection starts with a name lookup, and asyncio performs it by calling the blocking socket.getaddrinfo in the loop's default thread pool. That is invisible until the pool is busy — with file I/O, a synchronous SDK, asyncio.to_thread calls — at which point lookups queue behind unrelated work and every new connection waits. Measured on Python 3.14 on a 24-core machine, where the default executor has min(32, cpu_count + 4) = 28 workers: with the pool idle, 50 concurrent loop.getaddrinfo("localhost") calls took 1.4 ms at the median. With 28 one-second blocking jobs occupying the pool, the same lookups took 951 ms each — they waited for a thread. aiodns, which drives the c-ares library from the event loop without threads, answered in 1.82 ms idle and 1.69 ms with the pool full. Against real names through the system's caching resolver, first lookups took 0.3–20.7 ms either way. This guide keeps name resolution from becoming a hidden thread-pool dependency.

Prerequisites

1. See where asyncio resolves names

loop.getaddrinfo, which open_connection, create_connection and most HTTP clients call, is a thin wrapper:

# asyncio's implementation, in essence
async def getaddrinfo(self, host, port, *, family=0, type=0, proto=0, flags=0):
    return await self.run_in_executor(None, socket.getaddrinfo,
                                      host, port, family, type, proto, flags)

run_in_executor(None, ...) means the default ThreadPoolExecutor, shared with every asyncio.to_thread call and every run_in_executor(None, ...) in the process. A lookup is fast when a thread is free, and waits for one when none is. Nothing in an HTTP client's timeout configuration distinguishes "waiting for a DNS thread" from "waiting for the network", so the symptom appears as unexplained connect latency.

Verify: find every to_thread and run_in_executor(None, ...) in the codebase; each competes with name resolution for the same threads.

Lookup latency with the default executor idle and full 4 horizontal bars comparing loop.getaddrinfo, executor idle with the others. Lookup latency with the default executor idle and full loop.getaddrinfo, executor idle 1.4 ms loop.getaddrinfo, 28 workers busy 951 ms aiodns, executor idle 1.82 ms aiodns, 28 workers busy 1.69 ms 50 concurrent lookups of localhost; busy = 28 asyncio.to_thread(time.sleep, 1.0) jobs. Python 3.14, 24 cores. The thread-based resolver inherits every delay of the shared pool.

2. Resolve with aiodns

aiodns wraps c-ares, an asynchronous DNS library; its sockets are registered with the event loop, so a lookup is just another awaited I/O operation:

import socket
import aiodns


async def resolve(resolver: aiodns.DNSResolver, host: str, port: int) -> list[tuple[str, int]]:
    result = await resolver.getaddrinfo(host, port=port, type=socket.SOCK_STREAM)
    return [(node.addr[0].decode(), node.addr[1]) for node in result.nodes]   # addr[0] is bytes


async def main() -> None:
    resolver = aiodns.DNSResolver()                 # create once, inside the running loop
    print(await resolve(resolver, "www.python.org", 443))

Measured: 1.69–1.82 ms for local lookups whether or not the thread pool was busy, and 0.2–19.4 ms for real names through the system's resolver. c-ares reads /etc/hosts and /etc/resolv.conf itself, so it follows the same configuration as the system resolver for ordinary setups; it does not use NSS modules, so environments relying on LDAP or mDNS host lookups need the system resolver. One resolver per event loop is enough and should be reused: it holds its own sockets and, in recent c-ares versions, a query cache.

Verify: with the default executor deliberately saturated in a test, new connections still resolve in milliseconds.

3. Plug aiodns into HTTP clients

Clients make their own choice of resolver. aiohttp exposes it directly:

import aiohttp
from aiohttp.resolver import AsyncResolver

connector = aiohttp.TCPConnector(
    resolver=AsyncResolver(),        # aiodns-backed; requires the aiodns package
    ttl_dns_cache=300,               # aiohttp's own cache, in seconds
    limit=100,
)
session = aiohttp.ClientSession(connector=connector)

aiohttp's default ThreadedResolver uses loop.getaddrinfo and therefore the default executor; AsyncResolver does not. aiohttp also caches resolutions per connector (ttl_dns_cache, ten seconds by default), which matters more than the resolver for steady traffic to a few hosts and is covered in caching DNS lookups in async HTTP clients. httpx connects through anyio, whose asyncio backend calls loop.getaddrinfo — the same default executor; with httpx, the most effective protection is keeping connections alive so lookups happen rarely, and keeping the default executor free.

Verify: the client's configuration names its resolver explicitly, and DNS cache hit rates are visible in metrics or logs.

Which resolver each component uses A grid of 5 rows by 3 columns. Which resolver each component uses component resolver uses the default executor asyncio.open_connection / create_connection loop.getaddrinfo yes aiohttp TCPConnector (default) ThreadedResolver yes aiohttp TCPConnector(resolver=AsyncResolver()) aiodns / c-ares no httpx (via anyio) loop.getaddrinfo yes aiodns.DNSResolver directly c-ares on the event loop no Measured: a full executor made thread-based lookups wait 951 ms.

4. Or give blocking work its own executor

Sometimes the right fix is not to change the resolver but to stop unrelated blocking work from occupying the default pool. Give heavy blocking workloads a dedicated executor, leaving the default one for short tasks such as lookups:

from concurrent.futures import ThreadPoolExecutor

FILE_IO = ThreadPoolExecutor(max_workers=16, thread_name_prefix="file-io")
SDK = ThreadPoolExecutor(max_workers=32, thread_name_prefix="legacy-sdk")


async def save_report(path: str, data: bytes) -> None:
    loop = asyncio.get_running_loop()
    await loop.run_in_executor(FILE_IO, write_file, path, data)     # not the default pool


async def charge(card: str, amount: int) -> str:
    loop = asyncio.get_running_loop()
    return await loop.run_in_executor(SDK, legacy_client.charge, card, amount)

Separate pools are bulkheads: a slow SDK can exhaust its own threads without delaying name resolution or anything else that uses the default pool. The bulkhead idea, applied to dependencies rather than threads, is in bulkhead isolation with per-dependency semaphores. Raising the default pool's size with loop.set_default_executor is the blunt alternative; it delays the problem rather than isolating it.

Verify: under a load test that saturates the SDK pool, connection setup times to unrelated hosts stay flat.

5. Measure resolution time separately

To know whether DNS is the problem, time it separately from connecting. aiohttp's tracing hooks report resolution start and end per request:

from aiohttp import TraceConfig


async def on_dns_start(session, ctx, params):
    ctx.dns_start = asyncio.get_running_loop().time()


async def on_dns_end(session, ctx, params):
    DNS_SECONDS.observe(asyncio.get_running_loop().time() - ctx.dns_start)


trace = TraceConfig()
trace.on_dns_resolvehost_start.append(on_dns_start)
trace.on_dns_resolvehost_end.append(on_dns_end)
session = aiohttp.ClientSession(connector=connector, trace_configs=[trace])

A DNS-time histogram makes the executor problem diagnosable: lookups that are normally sub-millisecond but occasionally take hundreds of milliseconds, in step with spikes in to_thread usage, are waiting for threads, not for name servers. Tracing for HTTP clients more broadly is in instrumenting httpx and aiohttp with OpenTelemetry.

Verify: DNS time appears as its own metric, separate from connect and TLS time.

How should this service resolve names? A decision on What else uses the default thread pool with 4 outcomes. How should this service resolve names? What else uses the default thread pool? to_thread, file I/O, sync SDKs aiodns, or dedicated executors 1.7 ms vs 951 ms an aiohttp client AsyncResolver + ttl_dns_cache no threads, fewer lookups hosts from LDAP, mDNS, NSS system resolver, pool kept free c-ares skips NSS anything a DNS-time histogram see the stall Name resolution should never wait for an unrelated thread.

Verification

Name resolution is not a hidden bottleneck when:

  • Lookups do not share a saturated thread pool: aiodns is used, or heavy blocking work has its own executors.
  • HTTP clients name their resolver and cache resolutions.
  • DNS time is measured separately from connect and TLS time.
  • A test with the default executor saturated still connects in milliseconds.

Diagnostic Hook: chart DNS resolution time next to the number of busy threads in the default executor. Correlated spikes mean lookups are queueing for threads; flat DNS time with slow connects points elsewhere — the network, the TLS handshake or the server's accept queue.

Pitfalls & edge cases

  • to_thread floods. Measured: lookups waited 951 ms behind 28 blocking jobs.
  • Assuming the HTTP client resolves asynchronously. Most default to a thread.
  • c-ares and NSS. aiodns does not consult NSS modules such as LDAP or mDNS.
  • No DNS cache. Every new connection repeats the lookup; cache at the client.

Frequently Asked Questions

Is DNS resolution in asyncio non-blocking?

It does not block the event loop, but loop.getaddrinfo runs socket.getaddrinfo in the default thread pool, so it waits whenever that pool is busy: lookups took 951 ms in testing with all 28 workers occupied.

Should I use aiodns with asyncio?

When the default executor is shared with blocking work, yes: aiodns resolved names in about 1.7 ms whether or not the pool was busy. It does not use NSS modules, so keep the system resolver if you rely on LDAP or mDNS host lookups.

How do I make aiohttp resolve DNS asynchronously?

Pass TCPConnector(resolver=AsyncResolver()) with the aiodns package installed, and set ttl_dns_cache to cache results.

How big is asyncio's default thread pool?

min(32, os.cpu_count() + 4) workers — 28 on a 24-core machine — shared by to_thread, run_in_executor(None, ...) and getaddrinfo.