Wrapping Sync Cloud SDKs with asyncio.to_thread¶
Many cloud SDKs have no async client, or have one that lags behind the synchronous SDK in features. asyncio.to_thread lets async code call them without blocking the event loop: each call runs in the default thread pool while the loop serves everything else. The pool decides how many calls can run at once, and its default is min(32, os.cpu_count() + 4) threads — 28 on the 24-core test machine. Measured with boto3 fetching objects from a local S3-compatible server behind a proxy that added 100 ms per request: with the default executor, throughput was about 240 objects per second (237–241), regardless of botocore's own pool size; with a 64-thread executor it was 415. At 20 ms latency, boto3 in threads reached 445 objects per second while spending 2.47 ms of CPU per object, against 1.08 ms for the async client — threads contend for the GIL during signing and parsing. This guide wraps a sync SDK correctly and sizes the threads, the SDK's pool and the concurrency together.
Prerequisites¶
- Python 3.11+,
pip install boto3(or any synchronous SDK). - Thread offloading basics, from running blocking SDK calls with asyncio.to_thread.
- Thread safety of your SDK's objects: boto3 clients are thread-safe; sessions and resources are not.
1. Share one thread-safe client and call it through to_thread¶
Create the SDK client once, then wrap each blocking call:
import asyncio
import boto3
from botocore.config import Config
s3 = boto3.client("s3", config=Config(max_pool_connections=64)) # one client, thread-safe
async def get_bytes(bucket: str, key: str) -> bytes:
def call() -> bytes:
response = s3.get_object(Bucket=bucket, Key=key)
return response["Body"].read() # read inside the thread too
return await asyncio.to_thread(call)
Do the whole blocking operation inside the thread — including reading the streaming body, which is itself blocking I/O. Returning the response and reading the body on the event loop would block the loop for the transfer. boto3 documents clients as thread-safe and sessions as not, so create the client up front from one session rather than per thread or per call. For SDKs that are not thread-safe, give each worker thread its own client through threading.local.
Verify: with asyncio debug mode, no slow-callback warnings appear while many SDK calls run.
2. Size the executor for the concurrency you want¶
Every concurrent call occupies a thread for its full duration, including network latency. The default executor caps concurrency at min(32, cpu_count + 4):
import os
from concurrent.futures import ThreadPoolExecutor
SDK_CONCURRENCY = 64
async def main() -> None:
loop = asyncio.get_running_loop()
loop.set_default_executor(ThreadPoolExecutor(max_workers=SDK_CONCURRENCY,
thread_name_prefix="sdk"))
...
Measured at 100 ms latency: the default 28-thread executor gave about 240 objects per second — close to 28 threads ÷ 0.11 s per call — whether botocore's pool held 10 or 64 connections; 64 threads gave 415. Little's law gives the number of threads needed: target rate × call latency. The executor is shared with every other to_thread and run_in_executor call in the process, so a dedicated executor for the SDK keeps it from starving, or being starved by, other blocking work:
sdk_executor = ThreadPoolExecutor(max_workers=SDK_CONCURRENCY, thread_name_prefix="sdk")
async def get_bytes(bucket: str, key: str) -> bytes:
loop = asyncio.get_running_loop()
return await loop.run_in_executor(sdk_executor, fetch_sync, bucket, key)
Verify: with SDK_CONCURRENCY calls in flight, the executor has no queued work items and the SDK's connection pool is not exhausted.
3. Match the SDK's connection pool to the threads¶
Most SDKs keep their own HTTP connection pool, sized independently of your threads. boto3's default max_pool_connections is 10:
s3 = boto3.client("s3", config=Config(max_pool_connections=SDK_CONCURRENCY))
With more threads than pooled connections, urllib3 (under botocore) opens extra connections and discards them afterwards, logging "Connection pool is full, discarding connection" — throughput survives, as the 237 versus 241 measurement shows, but every extra request pays a new TCP and TLS handshake. Set the SDK's pool to the executor's size so each thread reuses a pooled connection. The same pattern applies to other SDKs: Google's clients take a transport or session with its own pool size, Azure's take a transport with connection limits.
Verify: no "pool is full" warnings in the logs at full concurrency, and connection setups per second stay near zero in steady state.
4. Expect GIL contention at high call rates¶
Threads help with waiting, not with Python work. SDKs do real CPU work per call — building and signing requests, parsing XML or JSON responses — and that work holds the GIL:
start_cpu, start = time.process_time(), time.perf_counter()
await asyncio.gather(*(get_bytes(bucket, k) for k in keys))
elapsed = time.perf_counter() - start
print(f"{len(keys) / elapsed:.0f}/s, {(time.process_time() - start_cpu) / len(keys) * 1000:.2f} ms CPU each")
# measured (20 ms latency, 32 threads): 445/s, 2.47 ms CPU per object
# same workload with aioboto3: 833/s, 1.08 ms CPU per object
The thread version spent more than twice the CPU per object for the same work, much of it in lock contention and context switching between 32 threads competing for one interpreter lock. That caps throughput per process and leaves less CPU for the rest of the application. When an async client exists and the call rate is high, it is the better choice, as compared in downloading many S3 objects concurrently. On free-threaded Python builds (3.13t and later) this contention largely disappears, provided the SDK and its dependencies support them.
Verify: CPU per call measured at your target rate; if it rises with thread count, you are in the contention regime.
5. Handle cancellation and timeouts around thread calls¶
A cancelled to_thread call stops your coroutine from waiting, but the thread keeps running the SDK call to completion. Timeouts belong inside the SDK call, where they can actually stop it:
s3 = boto3.client("s3", config=Config(
connect_timeout=5,
read_timeout=30, # bounds the blocking call in the thread
retries={"mode": "adaptive", "max_attempts": 5},
max_pool_connections=SDK_CONCURRENCY,
))
async def get_bytes(bucket: str, key: str, budget: float = 40.0) -> bytes:
async with asyncio.timeout(budget): # bounds the caller's wait, not the thread
return await asyncio.to_thread(fetch_sync, bucket, key)
asyncio.timeout around to_thread gives the caller a deadline, but the thread finishes its call regardless and occupies a pool slot until then; under a burst of timeouts, the pool fills with orphaned calls. Set the SDK's own connect and read timeouts below the async budget so blocked threads give up too. The general behaviour of timeouts on threads is covered in timing out blocking calls in threads.
Verify: after a burst of caller timeouts against a hanging endpoint, the executor's threads are free again within the SDK's read timeout.
Verification¶
A wrapped sync SDK behaves well when:
- One thread-safe client is shared, and whole operations, including body reads, run in the thread.
- The executor is sized from target rate × latency, ideally dedicated to the SDK.
- The SDK's connection pool matches the number of threads.
- SDK-level timeouts bound each call, since cancellation does not stop threads.
Diagnostic Hook: export the SDK executor's queue length (executor._work_queue.qsize() in a debug metric), calls in flight, and CPU per call. A growing queue with idle CPU means too few threads; CPU per call rising with thread count means GIL contention, the signal to move hot paths to an async client or more processes.
Pitfalls & edge cases¶
- The default executor for a busy SDK. Measured: capped near 240 calls per second at 100 ms latency.
- Reading the response body on the loop. It is blocking I/O; read it in the thread.
- SDK pool smaller than the thread count. Connections churn with "pool is full" warnings.
- Relying on cancellation to stop calls. Threads run to completion; set SDK timeouts.
Frequently Asked Questions¶
How do I call boto3 from asyncio without blocking?
Wrap each call, including reading the response body, in asyncio.to_thread, share one boto3 client, and size the thread pool and the client's max_pool_connections for your concurrency.
How many threads does asyncio.to_thread use?
The default executor has min(32, os.cpu_count() + 4) threads, 28 on a 24-core machine. In testing that capped boto3 at about 240 calls per second with 100 ms latency; 64 threads reached 415.
Is boto3 thread-safe?
boto3 clients are thread-safe and can be shared across threads; sessions and resources are not, so create clients from one session up front.
Why is boto3 in threads slower than aioboto3?
Signing and parsing hold the GIL, so threads contend. In testing boto3 in threads used 2.47 ms CPU per object against 1.08 ms for aioboto3, and reached 445 against 833 objects per second.
Related¶
- Cloud SDKs & Object Storage — up to the topic overview.
- Managing aioboto3 clients without leaking connections — the async alternative's lifecycle.
- Network I/O & Protocol Handling — the section overview.