Skip to content

Downloading Many S3 Objects Concurrently

Fetching thousands of objects one after another spends almost all its time waiting for round trips. Object stores handle concurrent requests well, so concurrency is the main lever — bounded, and matched to the client's connection pool. Measured with aioboto3 15.5 against a local S3-compatible server behind a proxy that added 20 ms of latency to each request, fetching 600 objects of 100 KB: 37 objects per second one at a time, 293 with 8 concurrent, 833 with 32, and 801 with 64 — where the test server itself saturated. boto3 running in 32 threads reached 445 per second and spent 2.47 ms of CPU per object against 1.08 ms for aioboto3. This guide lists keys as a stream, fans downloads out with a bound, and handles failures per key.

Prerequisites

1. Bound concurrency and match the pool

A semaphore sets the number of downloads in flight; the client's pool must be at least that large, or extra requests wait inside the pool:

from aiobotocore.config import AioConfig

CONCURRENCY = 32
config = AioConfig(max_pool_connections=CONCURRENCY)


async def fetch_all(s3, bucket: str, keys: list[str]) -> dict[str, bytes]:
    slots = asyncio.Semaphore(CONCURRENCY)

    async def one(key: str) -> tuple[str, bytes]:
        async with slots:
            response = await s3.get_object(Bucket=bucket, Key=key)
            async with response["Body"] as body:
                return key, await body.read()

    return dict(await asyncio.gather(*(one(k) for k in keys)))

Measured with 20 ms of latency: throughput rose almost linearly from 37 to 293 objects per second between concurrency 1 and 8, and to 833 at 32, where the local server became the limit. Against a real object store, the ceiling is usually your bandwidth, CPU, or the provider's per-prefix request rate rather than the server. Start at 16–32 and measure; more is not better once throughput stops rising.

Verify: throughput at your chosen concurrency is within a few percent of the next doubling.

Objects per second by concurrency, 20 ms added latency 5 horizontal bars comparing aioboto3, concurrency 1 with the others. Objects per second by concurrency, 20 ms added latency aioboto3, concurrency 1 37/s aioboto3, concurrency 8 293/s aioboto3, concurrency 32 833/s aioboto3, concurrency 64 801/s (server limit) boto3, 32 threads 445/s 600 objects of 100 KB, local S3-compatible server behind a proxy adding 20 ms per request. Concurrency multiplies throughput until something else saturates.

2. Stream keys from the listing instead of collecting them

For large prefixes, listing every key before downloading anything wastes time and memory. Feed keys from the paginator into a bounded queue that workers consume:

async def download_prefix(s3, bucket: str, prefix: str, sink, workers: int = 32) -> Stats:
    queue: asyncio.Queue[str | None] = asyncio.Queue(maxsize=workers * 4)
    stats = Stats()

    async def producer() -> None:
        paginator = s3.get_paginator("list_objects_v2")
        async for page in paginator.paginate(Bucket=bucket, Prefix=prefix):
            for obj in page.get("Contents", []):
                await queue.put(obj["Key"])                # waits when workers are behind
        for _ in range(workers):
            await queue.put(None)                          # one stop signal per worker

    async def worker() -> None:
        while (key := await queue.get()) is not None:
            await download_one(s3, bucket, key, sink, stats)

    async with asyncio.TaskGroup() as tg:
        tg.create_task(producer())
        for _ in range(workers):
            tg.create_task(worker())
    return stats

Downloads start after the first page of 1,000 keys instead of after the last, and memory holds at most a few pages of keys regardless of the prefix size. Listing is sequential — each page needs the previous page's continuation token — so for very large buckets, list several known sub-prefixes concurrently and feed the same queue.

Verify: the first download starts within one listing round trip, and memory stays flat for prefixes of millions of keys.

3. Handle failures per key

In a batch of thousands, some downloads fail: a key deleted since listing, a throttling response, a timeout. One failure should not cancel the batch:

from botocore.exceptions import ClientError


async def download_one(s3, bucket: str, key: str, sink, stats: "Stats") -> None:
    for attempt in range(3):
        try:
            response = await s3.get_object(Bucket=bucket, Key=key)
            async with response["Body"] as body:
                await sink(key, await body.read())
            stats.ok += 1
            return
        except ClientError as exc:
            code = exc.response["Error"]["Code"]
            if code in ("NoSuchKey", "AccessDenied"):
                stats.failed[key] = code                    # permanent: record, move on
                return
            if code in ("SlowDown", "503", "RequestTimeout") and attempt < 2:
                await asyncio.sleep(0.5 * 2 ** attempt)     # throttled: back off
                continue
            stats.failed[key] = code
            return
        except (TimeoutError, aiohttp.ClientError) as exc:
            if attempt == 2:
                stats.failed[key] = type(exc).__name__
                return

The client already retries some errors internally (with retries={"mode": "adaptive"} it also slows its request rate when throttled); the loop above adds per-key classification so the batch finishes with a report instead of an exception. Write the failures to a file of keys, so a second pass can retry exactly those — the same approach as in downloading many URLs concurrently with progress.

Verify: a batch containing deleted keys completes, reports them as NoSuchKey, and downloads everything else.

Downloading a prefix A flow of 5 stages. Downloading a prefix paginate listing pages of 1,000 keys bounded queue backpressure on listing N workers pool >= N per-key errors record or retry sink + failure list rerun failures Listing, downloading and error handling overlap; memory stays bounded.

4. Keep CPU in mind at high rates

At hundreds of objects per second, request signing and response parsing become significant CPU work. Measured with 20 ms latency and 32 concurrent downloads: aioboto3 used 1.08 ms of CPU per object; boto3 in 32 threads used 2.47 ms per object and topped out at 445 objects per second, because botocore's work holds the GIL and the threads contend for it:

start_cpu, start = time.process_time(), time.perf_counter()
stats = await download_prefix(s3, bucket, prefix, sink)
cpu_ms_per_object = (time.process_time() - start_cpu) / stats.ok * 1000
log.info("%.0f objects/s, %.2f ms CPU per object",
         stats.ok / (time.perf_counter() - start), cpu_ms_per_object)

When one process's CPU becomes the limit, split the key space across processes — by prefix or by a hash of the key — each with its own client and concurrency. Small objects make this matter sooner, because the per-request cost is fixed while the bytes are few.

Verify: at target throughput, process CPU stays below one core per process, or the work is split across processes.

5. Write results safely

Where downloaded objects go determines what a crash leaves behind. For local files, write to a temporary name and rename; for databases, upsert by key:

async def file_sink(key: str, data: bytes, root: str) -> None:
    path = os.path.join(root, key)
    os.makedirs(os.path.dirname(path), exist_ok=True)
    await asyncio.to_thread(write_atomic, path, data)     # temp file + fsync + os.replace


def already_have(root: str, key: str, size: int) -> bool:
    path = os.path.join(root, key)
    return os.path.exists(path) and os.path.getsize(path) == size

Checking for an existing file of the right size before downloading makes a rerun after a crash skip completed keys, so the second pass costs only the remainder. The atomic write pattern is in writing files atomically from async code; for objects larger than a few megabytes, stream them to disk instead of reading them into memory, as in streaming object downloads to disk without buffering.

Verify: kill a download run halfway and rerun it; completed files are skipped and no partial file exists under a final name.

How should this batch of objects be downloaded? A decision on What does the batch look like with 4 outcomes. How should this batch of objects be downloaded? What does the batch look like? known list, thousands semaphore + gather pool = concurrency huge prefix paginator -> queue -> workers flat memory CPU-bound rate split keys across processes one client each large objects stream to disk not body.read() Bound the fan-out, stream the listing, and isolate failures per key.

Verification

Bulk downloads work well when:

  • Concurrency is bounded and the client pool is at least as large.
  • Keys stream from the paginator into a bounded queue.
  • Failures are classified per key, retried when transient, and reported.
  • Results are written atomically and reruns skip completed keys.

Diagnostic Hook: chart objects per second, CPU per object, throttling errors and pool waits. Throughput that stops rising as concurrency increases while CPU per object stays flat points at the server, network or provider limits; CPU per object rising with concurrency points at GIL contention in a thread-based client.

Pitfalls & edge cases

  • Concurrency above the pool size. Extra requests wait inside the client.
  • Listing everything first. Delays the first download and holds every key in memory.
  • One failure cancelling the batch. Classify and record per key instead.
  • Sync SDK threads at high rates. Measured 2.47 ms CPU per object against 1.08 ms.

Frequently Asked Questions

How do I download many S3 objects in parallel with Python asyncio?

Share one aioboto3 client with max_pool_connections at least your concurrency, bound downloads with an asyncio.Semaphore or a fixed set of workers, and read each body inside async with. In testing at 20 ms latency, concurrency 32 gave 833 objects per second against 37 one at a time.

What concurrency should I use for S3 downloads?

Start around 16 to 32 and measure: throughput rises until bandwidth, CPU, the provider's request limits or the server saturate, then flattens.

Is aioboto3 faster than boto3 with threads?

For many small objects, yes: 833 versus 445 objects per second in testing, with less than half the CPU per object, because botocore's work in threads contends for the GIL.

How do I list millions of S3 keys without running out of memory?

Iterate the list_objects_v2 paginator with async for and push keys into a bounded queue consumed by download workers, instead of collecting all keys into a list first.