Skip to content

Offering Sync and Async Versions of One API

A client library that only offers async def methods loses every Flask app, script and notebook; one that only offers blocking methods loses every asyncio service. Shipping both is common — httpx, redis-py, the OpenAI and Anthropic SDKs and most database drivers do it — and the ways of doing it differ enormously in cost and correctness. The tempting one, wrapping each async method in asyncio.run(), cost 45.6 µs per call in a microbenchmark and fails outright inside a running loop with RuntimeError: asyncio.run() cannot be called from a running event loop. This guide compares the four workable designs and shows when each fits.

Prerequisites

1. Rule out asyncio.run per call

The one-line wrapper is the first thing everyone writes:

class Client:
    def get(self, key: str) -> bytes:
        return asyncio.run(self._aclient.get(key))     # new loop every call

It has three problems, in increasing order of severity. Each call creates and closes an event loop — 45.6 µs per call against 0.05 µs for a plain function call in the same benchmark. Any async resource the client holds, such as a connection pool, is bound to the loop it was created on, so it cannot survive from one asyncio.run() to the next; the "client" reconnects on every call. And it raises RuntimeError when called from code that is already inside an event loop — which includes Jupyter, most async test runners, and any async framework calling into "sync" helper code.

asyncio.Runner (3.11+) fixes the first two by keeping one loop alive across calls: 8.1 µs per call in the same benchmark, and resources survive. It does not fix the third.

Verify: call the sync wrapper from inside an async def and confirm you get the RuntimeError — then decide whether your users ever do that.

Per-call overhead of sync wrappers around one async call 4 horizontal bars comparing asyncio.run per call with the others. Per-call overhead of sync wrappers around one async call asyncio.run per call 45.6 µs background loop thread 25.2 µs asyncio.Runner, one loop 8.1 µs plain sync function 0.05 µs Python 3.14, Linux; the wrapped coroutine does a single await asyncio.sleep(0). Overhead is small next to network I/O; correctness inside a running loop is what separates the options.

2. Run a private loop in a background thread

The wrapper that works everywhere — including inside a running loop — owns a dedicated event loop on a daemon thread and submits coroutines to it:

import asyncio
import threading


class _LoopThread:
    def __init__(self) -> None:
        self.loop = asyncio.new_event_loop()
        self._thread = threading.Thread(target=self.loop.run_forever, name="client-loop", daemon=True)
        self._thread.start()

    def run(self, coro, timeout: float | None = None):
        return asyncio.run_coroutine_threadsafe(coro, self.loop).result(timeout)

    def close(self) -> None:
        self.loop.call_soon_threadsafe(self.loop.stop)
        self._thread.join()
        self.loop.close()


class Client:
    def __init__(self) -> None:
        self._lt = _LoopThread()
        self._async = self._lt.run(_make_async_client())   # pool lives on the private loop

    def get(self, key: str) -> bytes:
        return self._lt.run(self._async.get(key), timeout=30)

Calls cost 25.2 µs each — a thread hop in each direction — and the async client's pool lives for the life of the sync client. Because the caller blocks on a concurrent.futures.Future rather than running a loop, it works from plain threads, Flask views, and even from inside another event loop (where it blocks that loop, which is the caller's problem to avoid with to_thread).

Its cost is architectural: every sync client carries a thread, and exceptions, cancellation and timeouts cross a thread boundary. The deeper version is in running an event loop in a background thread.

Verify: call Client().get() from inside a running loop via asyncio.to_thread, and from a plain thread; both must work.

3. Generate the sync code from the async code

The approach httpcore uses is to write the async implementation once and mechanically generate a sync copy by rewriting async def to def, removing await, and renaming async classes. The unasync tool does exactly this at build time:

# setup / build step
import unasync

unasync.unasync_files(
    ["src/mylib/_async/client.py", "src/mylib/_async/pool.py"],
    rules=[unasync.Rule(
        fromdir="/_async/",
        todir="/_sync/",
        additional_replacements={"AsyncClient": "Client", "anyio": "_sync_backend"},
    )],
)
# src/mylib/_async/client.py — the only file a human edits
class AsyncClient:
    async def get(self, key: str) -> bytes:
        async with self._pool.connection() as conn:
            await conn.send(encode_get(key))
            return await conn.read_reply()

The generated _sync/client.py has the same logic with blocking calls, so the sync API has zero event-loop overhead and no hidden thread. The price is discipline: the async source may only use constructs that have a mechanical sync equivalent, and the I/O primitives underneath (conn.send) need real sync and async implementations. Concurrency features with no sync counterpart — TaskGroup, asyncio.wait — cannot appear in the shared code.

Verify: generate, then run the same test suite against both the async and generated sync clients.

One source, two generated APIs A flow of 4 stages. One source, two generated APIs _async/client.py the only edited file unasync at build strip async, await _sync/client.py generated, committed one test suite runs against both Code generation removes runtime overhead entirely, at the cost of restricting what the async source may use.

4. Share a sans-I/O core with two thin drivers

When the library speaks a wire protocol, the cleanest split is to put all protocol logic in a sans-I/O core and write two small drivers:

class _Core:                                    # pure: no I/O
    def build_request(self, key: str) -> bytes: ...
    def parse_reply(self, data: bytes) -> Reply | None: ...


class Client:
    def get(self, key: str) -> bytes:
        self._sock.sendall(self._core.build_request(key))
        while (reply := self._core.parse_reply(self._sock.recv(65536))) is None:
            pass
        return reply.value


class AsyncClient:
    async def get(self, key: str) -> bytes:
        self._writer.write(self._core.build_request(key))
        await self._writer.drain()
        while (reply := self._core.parse_reply(await self._reader.read(65536))) is None:
            pass
        return reply.value

The drivers are short and differ only in their I/O calls; everything that is hard to get right — framing, state, errors, limits — is shared and tested once. This is the design behind h11/h2 (used by both httpx sync and async clients) and wsproto.

Verify: the core module has no async def and no socket imports, and both drivers pass the same protocol test vectors.

5. Choose by what the library does

The options trade differently, and the right one depends on how much of the library is I/O and how much is logic:

Approach Sync overhead Works inside a running loop Maintenance
asyncio.run per call 45.6 µs, no pooling no trivial, but wrong
asyncio.Runner 8.1 µs no trivial
background loop thread 25.2 µs yes, blocks caller one wrapper class
code generation none yes build step, restricted async subset
sans-I/O core none yes two drivers per I/O call site

For an SDK wrapping an HTTP API, the background loop thread or code generation over httpx's own sync and async clients are the pragmatic choices. For a protocol library, sans-I/O pays for itself in testability. asyncio.run per call is only acceptable for scripts that make a handful of calls and never run inside a loop.

Verify: benchmark your sync API against the raw async API under the same load; the overhead should be within the budget you set for the library.

Which dual-API design fits this library? A decision on What does the library mostly consist of with 3 outcomes. Which dual-API design fits this library? What does the library mostly consist of? a wire protocol sans-I/O core two thin drivers logic over a dual client generate with unasync zero sync overhead calls into async-only deps background loop thread one wrapper Never ship asyncio.run per call as the sync API of a library that holds connections.

Verification

A dual API is shipped correctly when:

  • Both APIs pass one shared test suite, parametrised over sync and async clients.
  • The sync API works inside a running loop (via to_thread) or documents clearly that it does not.
  • Connection pools persist across sync calls — the server sees one connection reused, not one per call.
  • Closing the sync client stops any loop thread and closes pooled connections.

Diagnostic Hook: expose the same metrics from both clients — requests, errors, pool in-use — with a api="sync" or api="async" label. A sync client whose connection count tracks its request count instead of staying flat is creating a loop per call; a sync client with a loop thread whose calls take much longer than the async equivalent under load is contending on that one thread and needs either a pool of loop threads or the generated design.

Pitfalls & edge cases

  • Sharing one async client between the private loop and the caller's loop. Async objects are bound to the loop that created them; never hand the sync wrapper's internal client to async callers.
  • Calling the sync API from async code without to_thread. It blocks the caller's loop for the full duration of the call.
  • Forgetting close() on a loop-thread client. The daemon thread dies at exit without closing connections, producing unclosed-resource warnings.
  • Async-only constructs in code fed to unasync. TaskGroup and asyncio.gather have no mechanical sync equivalent; the generated code will not run.

Frequently Asked Questions

How do I provide both sync and async versions of a Python API?

Write the logic once and expose it twice: generate the sync code from the async source with unasync, share a sans-I/O core between two thin drivers, or wrap the async client with a private event loop on a background thread. Avoid wrapping each call in asyncio.run.

Why not just call asyncio.run inside the sync method?

It creates a new event loop per call, so pooled connections cannot be reused, costs about 46 µs per call, and raises RuntimeError whenever the caller is already inside a running event loop, such as Jupyter or an async test.

How does httpx support sync and async?

httpx has separate Client and AsyncClient classes over the httpcore transport, which maintains its sync code by generating it from the async implementation, and both use the sans-I/O h11 and h2 libraries for the HTTP protocol itself.

Is a background event loop thread safe to use in a library?

Yes, if the library owns the loop entirely: create it privately, submit work with run_coroutine_threadsafe, and stop and join the thread on close. Never expose objects bound to that loop to callers on other loops.