Skip to content

Cloud SDKs & Object Storage with asyncio

Object storage is where most services keep files, exports, model artefacts and backups, and its APIs are plain HTTP — a natural fit for asyncio. The official SDKs, though, are mostly synchronous (boto3, the Google and Azure clients), and their async variants (aioboto3 and aiobotocore for AWS, async clients for Google and Azure) bring their own lifecycles and limits. Measured with aioboto3 15.5 against a local S3-compatible server: creating a client per request cost 25.9 ms per operation against 2.1 ms with a shared client; the default connection pool of 10 capped concurrent large downloads at about 10 no matter how many tasks were waiting; reading a 256 MiB object with await body.read() peaked at 577 MiB of memory while streaming it to disk in 1 MiB chunks peaked at 72 MiB and was faster. With 20 ms of simulated latency, concurrent downloads scaled from 37 objects per second at concurrency 1 to 833 at 32, while boto3 in threads topped out near 445 and spent 2.47 ms of CPU per object against 1.08 ms for aioboto3.

This section covers using object storage and cloud SDKs from asyncio without leaking connections, buffering whole objects, or capping throughput by accident. The parent section, Network I/O & Protocol Handling, covers the HTTP clients and pools these SDKs are built on.

Scope of this section:

  • Client and session lifecycles in aioboto3 and similar async SDKs.
  • Connection pool sizing and the bodies that hold connections.
  • Multipart uploads for large files, and cleaning up after failures.
  • Concurrent downloads of many objects, and streaming large ones to disk.
  • Using synchronous SDKs from asyncio through threads, and what that costs.

Architectural principles

  • One client per process, created at startup. SDK clients carry connection pools, credentials and endpoint metadata; creating them per request multiplies setup costs.
  • The pool is a concurrency limit. max_pool_connections caps simultaneous requests whether or not you meant it to; size it with the concurrency you intend.
  • Bodies hold connections. A response body not read to the end or closed keeps its connection out of the pool.
  • Stream large objects. Uploads go in parts, downloads in chunks; memory should depend on chunk size and concurrency, never on object size.
  • Clean up partial work. Incomplete multipart uploads are stored, billed and invisible; abort them on failure and expire them by policy.
The layers between your code and the object store 5 stacked layers. The layers between your code and the object store application tasks bounded by a semaphore one async client created at startup, closed at shutdown connection pool max_pool_connections (default 10) response bodies hold a connection until read/closed object store HTTP API S3, GCS, Azure Blob Each layer has a limit; the tightest one sets your throughput.

Execution model: async SDKs over aiohttp

aioboto3 wraps aiobotocore, which runs botocore's request building, signing and parsing and replaces its HTTP layer with aiohttp. That has two consequences. First, the client is an async context manager that owns an aiohttp connector: it must be entered once and exited at shutdown, or its connections leak and aiohttp logs "Unclosed connector". Second, every response body is an aiohttp stream: the connection returns to the pool only when the body has been read to the end or closed, so a forgotten body is a leaked connection until garbage collection. Tested with a pool of 2, two unread 4 MiB bodies made a third request wait until its 3 s timeout; closing them let it complete in 35 ms.

import aioboto3
from aiobotocore.config import AioConfig

session = aioboto3.Session()


@asynccontextmanager
async def lifespan(app):
    async with session.client("s3", config=AioConfig(max_pool_connections=64)) as s3:
        app.state.s3 = s3                       # one client for the process
        yield                                   # closed (and its pool) at shutdown


async def read_object(s3, bucket: str, key: str) -> bytes:
    response = await s3.get_object(Bucket=bucket, Key=key)
    async with response["Body"] as body:        # releases the connection when done
        return await body.read()

Synchronous SDKs called through asyncio.to_thread work too, with different limits: the default thread pool has min(32, cpu_count + 4) workers, and botocore's own work — signing, XML parsing — holds the GIL, so threads contend for it. Details are in wrapping sync cloud SDKs with asyncio.to_thread.

Cost per get_object by client lifecycle 2 horizontal bars comparing client per request with the others. Cost per get_object by client lifecycle client per request 25.9 ms/op shared client 2.1 ms/op aioboto3 15.5 against a local S3-compatible server, 10 KB objects. Client creation, not the request, dominated the per-request pattern.

Pattern catalogue

Share one client and size its pool

Create the client once and set max_pool_connections to the number of concurrent requests you intend. Measured with 50 concurrent 4 MiB downloads: the default pool of 10 allowed about 10 bodies open at once (peak 11) and took 0.78 s; a pool of 50 allowed 50 and took 0.58 s. See managing aioboto3 clients without leaking connections.

Upload large files in parts

A single put_object of a 256 MiB file read the file into memory and peaked at 579 MiB. Multipart upload sends fixed-size parts, several at a time: manual multipart with 8 MiB parts peaked at 91 MiB with one part in flight and 204 MiB with eight, and the managed upload_fileobj with 8 MiB parts and 8 concurrent reached 205 MiB/s at a 156 MiB peak. Failed uploads must be aborted — tested, an incomplete upload stayed listed until abort_multipart_upload. See uploading large files to S3 with async multipart.

from boto3.s3.transfer import TransferConfig

with open(path, "rb") as f:
    await s3.upload_fileobj(f, bucket, key,
                            Config=TransferConfig(multipart_chunksize=8 * 2**20, max_concurrency=8))

Download many objects with bounded concurrency

Object stores reward concurrency because each request mostly waits on latency. With 20 ms of added latency, aioboto3 fetched 37 objects per second one at a time, 293 at concurrency 8 and 833 at 32, where the test server saturated. Bound concurrency with a semaphore matched to the pool. See downloading many S3 objects concurrently.

Stream large downloads to disk

Read bodies in chunks and write them as they arrive. Measured on a 256 MiB object: await body.read() then write peaked at 577 MiB and ran at 303 MiB/s; iter_chunks(1 MiB) written through a thread peaked at 72 MiB and ran at 491 MiB/s; the managed download_file ran at 501 MiB/s and peaked at 189 MiB. See streaming object downloads to disk without buffering.

response = await s3.get_object(Bucket=bucket, Key=key)
with open(tmp_path, "wb") as f:
    async for chunk in response["Body"].iter_chunks(1024 * 1024):
        await asyncio.to_thread(f.write, chunk)

Run synchronous SDKs in threads, sized deliberately

When no async client exists, asyncio.to_thread keeps the loop free, but throughput is capped by the thread pool and by the GIL. With 100 ms latency, the default 28-thread executor delivered about 240 objects per second; 64 threads delivered 415. With 20 ms latency, boto3 in 32 threads topped out at 445 per second using 2.47 ms of CPU per object, against 1.08 ms for aioboto3.

Measured costs and limits, at a glance A grid of 5 rows by 3 columns. Measured costs and limits, at a glance behaviour measured fix client per request 25.9 vs 2.1 ms/op one client per process default pool (10) ~10 bodies open at once max_pool_connections await body.read() of 256 MiB 577 MiB peak iter_chunks: 72 MiB single put of 256 MiB 579 MiB peak multipart parts boto3 in threads 2.47 ms CPU/object aioboto3: 1.08 ms Local S3-compatible server; latency simulated with a delaying proxy where noted.

Choosing a client approach

Which way should this service talk to object storage? A decision on What does the workload look like with 4 outcomes. Which way should this service talk to object storage? What does the workload look like? many concurrent ops async SDK, one client pool = concurrency occasional calls, or no async SDK sync SDK + to_thread sized executor few huge files managed transfer (upload_fileobj) parts in parallel presigned URLs plain httpx/aiohttp no SDK in the hot path The async SDK wins on throughput and CPU; threads win on availability.

Async SDKs are the default for services that do many storage operations concurrently: they used less than half the CPU per object in testing and scale with the event loop. Synchronous SDKs through threads are fine for occasional calls — a nightly export, an admin action — and are the only option for services whose SDK has no async client. Presigned URLs move transfers out of the SDK entirely: the service signs a URL, and clients or a plain async HTTP client do the transfer, which is useful for very large objects and for offloading bandwidth to end users.

Resource boundaries

  • Connections: max_pool_connections per client (default 10 in botocore and aiobotocore); every open body holds one.
  • Memory: chunk size × concurrency for streaming; part size × concurrent parts for multipart uploads (8 MiB × 8 is about 64 MiB of buffers plus overhead).
  • Threads: for sync SDKs, the executor size caps concurrency; the default is min(32, cpu_count + 4).
  • Provider limits: request rates per prefix, part counts (10,000 per S3 upload) and part sizes (at least 5 MiB except the last — tested, 1 KiB parts failed with EntityTooSmall).
  • Cost: requests are billed individually; listing and many small objects can cost more than bandwidth.

Integrated production example

A small storage service combining the patterns: one client from the lifespan, bounded concurrent downloads streamed to disk, and multipart uploads through the managed transfer:

import asyncio
import os
from contextlib import asynccontextmanager

import aioboto3
from aiobotocore.config import AioConfig
from boto3.s3.transfer import TransferConfig

MAX_CONCURRENCY = 32
UPLOAD_CONFIG = TransferConfig(multipart_threshold=16 * 2**20, multipart_chunksize=8 * 2**20,
                               max_concurrency=8)


class Storage:
    def __init__(self, s3, bucket: str) -> None:
        self.s3, self.bucket = s3, bucket
        self.slots = asyncio.Semaphore(MAX_CONCURRENCY)

    async def download(self, key: str, path: str) -> int:
        tmp = path + ".part"
        async with self.slots:
            response = await self.s3.get_object(Bucket=self.bucket, Key=key)
            size = 0
            with open(tmp, "wb") as f:
                async for chunk in response["Body"].iter_chunks(1024 * 1024):
                    await asyncio.to_thread(f.write, chunk)
                    size += len(chunk)
        os.replace(tmp, path)
        return size

    async def download_many(self, keys: list[str], directory: str) -> dict[str, int | BaseException]:
        async def one(key: str):
            return await self.download(key, os.path.join(directory, key.replace("/", "_")))
        results = await asyncio.gather(*(one(k) for k in keys), return_exceptions=True)
        return dict(zip(keys, results))

    async def upload(self, path: str, key: str) -> None:
        with open(path, "rb") as f:
            await self.s3.upload_fileobj(f, self.bucket, key, Config=UPLOAD_CONFIG)


@asynccontextmanager
async def storage_lifespan(app):
    session = aioboto3.Session()
    config = AioConfig(max_pool_connections=MAX_CONCURRENCY, retries={"mode": "adaptive", "max_attempts": 5})
    async with session.client("s3", config=config) as s3:
        app.state.storage = Storage(s3, os.environ["BUCKET"])
        yield

The semaphore and the pool share one number, so tasks wait on the semaphore rather than inside the pool where they would count against request timeouts. Downloads go to a temporary name and are renamed when complete, as in writing files atomically from async code. The adaptive retry mode backs off client-side when the provider throttles.

Listing objects and working with many small ones

Listing is paginated — S3 returns at most 1,000 keys per list_objects_v2 call — and aiobotocore's paginators are async iterators, so a listing can feed a download pipeline as pages arrive instead of collecting millions of keys first:

async def iter_keys(s3, bucket: str, prefix: str):
    paginator = s3.get_paginator("list_objects_v2")
    async for page in paginator.paginate(Bucket=bucket, Prefix=prefix):
        for obj in page.get("Contents", []):
            yield obj["Key"], obj["Size"]

Each page is one request, and pages arrive in sequence, so listing a very large bucket is slow on its own; when prefixes are known (dates, tenants, shards), list several prefixes concurrently and merge the streams. For work over many small objects, request count dominates both latency and cost: batch deletions with delete_objects (up to 1,000 keys per call), avoid a head_object before every get_object when the get itself returns the metadata, and consider packing small items into larger objects when they are always read together. A thousand 1 KB objects cost a thousand requests to write and another thousand to read; one 1 MB object costs one of each. For keys that are written and read at high rates, spread them across prefixes, because providers scale request capacity per prefix.

Credentials, regions and retries

Async clients resolve credentials the same way their synchronous counterparts do — environment variables, shared config files, container and instance metadata endpoints, web-identity tokens — but the resolution happens when the client is created and, for temporary credentials, again when they near expiry. Two practical consequences follow. Creating the client at startup surfaces credential problems immediately, before traffic arrives, instead of on the first request. And a long-lived client refreshes temporary credentials itself; there is no need to recreate it periodically, which would throw away its connection pool.

Set the region explicitly rather than relying on the environment, so a misconfigured deployment fails fast instead of signing requests for the wrong region. For S3-compatible stores — MinIO, SeaweedFS, Ceph, Cloudflare R2 — pass endpoint_url; the API calls and every pattern in this section are the same. Configure retries on the client (retries={"mode": "adaptive"} in botocore's config): adaptive mode adds client-side rate limiting when the service throttles, which is better than layering your own retry loop on top and multiplying attempts.

Other providers

Google Cloud Storage and Azure Blob Storage have their own async clients — the gcloud-aio-storage package and azure.storage.blob.aio respectively — and the same principles carry over directly: one client per process, explicit session or transport lifecycles, connection limits set on the underlying HTTP session, streamed reads and chunked or block uploads for large objects. Their equivalent of multipart upload is resumable uploads (GCS) and staged blocks committed together (Azure). Where a provider only offers a synchronous SDK, the thread-based approach applies with the same sizing rules. For multi-cloud code, a thin storage interface — get, put, stream, delete, list — keeps provider details in one module.

Testing against object storage

Run tests against a local S3-compatible server in a container rather than mocking the SDK: the measurements in this section were taken against one, and it exercises real HTTP, real streaming bodies and real multipart rules — including the 5 MiB minimum part size, which a mock would happily ignore. Create a uniquely named bucket per test session, and clean up incomplete multipart uploads in teardown so failures do not accumulate between runs.

Diagnostic hook callout

Diagnostic Hook: export, per client, requests in flight, time spent waiting for a pooled connection, bytes transferred, and errors by code (throttling, timeouts, 5xx). Requests in flight pinned at max_pool_connections with rising pool waits means the pool is the limit; throttling errors (SlowDown, 503) mean the provider is the limit and concurrency should drop; process memory that tracks object sizes means something is reading whole bodies.

Failure modes

  • A client per request. Measured 25.9 ms of overhead per operation.
  • Default pool size. Concurrency silently capped at about 10.
  • Unread or unclosed bodies. Connections leak until requests time out waiting for the pool.
  • Whole-object reads and puts. Memory proportional to object size: 577–579 MiB for 256 MiB.
  • Orphaned multipart uploads. Stored and billed until aborted or expired.

Frequently Asked Questions

Should I use aioboto3 or boto3 with asyncio?

For services that make many concurrent storage calls, aioboto3: in testing it used 1.08 ms of CPU per object against 2.47 ms for boto3 in threads and scaled further. For occasional calls, boto3 through asyncio.to_thread is simpler and fine.

How many concurrent S3 requests can aioboto3 make?

As many as max_pool_connections allows, which defaults to 10. Raise it to your intended concurrency and bound tasks with a semaphore of the same size.

How do I download a large S3 object without loading it into memory?

Iterate the response body with iter_chunks and write each chunk to a file. A 256 MiB object peaked at 72 MiB of process memory that way, against 577 MiB with body.read().

What happens to a failed multipart upload?

Its uploaded parts stay stored, invisible as an object but still billed, until you call abort_multipart_upload or a lifecycle rule expires incomplete uploads.

Why do my S3 requests hang under load?

Usually every pooled connection is held by an unread response body or the pool is too small for the concurrency. Read or close every body and size max_pool_connections.