Skip to content

Listing Objects with Async Paginators

Listing a bucket is a sequence of ListObjectsV2 calls, each returning at most 1,000 keys and a token for the next page, so a single paginator is inherently sequential — and on large buckets that sequence is slow. The obvious fix, listing several prefixes concurrently, works only up to a point that has nothing to do with the network. Measured with aioboto3 15.5 against SeaweedFS holding 100,000 objects under 16 prefixes, through a proxy adding 30 ms per request: one paginator took 9.07–9.28 s, 100 pages at about 93 ms each. Sixteen paginators running concurrently, one per prefix, took 4.26–4.31 s — not the sixteenfold gain the round trips suggest, because the client spent 4.2 s of CPU parsing responses: 42–44 ms per 1,000-key page in botocore. The event loop was saturated, with lag reaching 1.2 s. Spreading the prefixes over processes took 2.22 s with two, 1.32 s with four and 0.99 s with eight, with the parent's loop lag at 2–6 ms. Streaming the keys instead of collecting them kept memory flat — +3 MiB against +47 MiB for 100,000 keys — and resuming an interrupted listing with StartAfter returned every key exactly once. This guide covers listing large buckets from async code.

Prerequisites

1. Stream pages instead of collecting keys

The async paginator yields one page at a time. Process each page as it arrives rather than accumulating every key first:

async def iter_keys(s3, bucket: str, prefix: str = ""):
    paginator = s3.get_paginator("list_objects_v2")
    async for page in paginator.paginate(Bucket=bucket, Prefix=prefix):
        for obj in page.get("Contents", []):
            yield obj["Key"], obj["Size"]

Measured over 100,000 keys: streaming raised peak RSS by 3 MiB; collecting the Contents dictionaries into a list raised it by 47 MiB — about 470 bytes per object, which is 4.7 GB at ten million. Streaming also lets downstream work start on the first page instead of after the last; the pattern of feeding keys into a bounded download stage is shown in downloading many S3 objects concurrently. Note that Contents is absent, not empty, on a page with no keys, hence page.get("Contents", []).

Verify: memory during a full listing is independent of the number of keys.

2. Measure where a page's time goes

Before parallelising, find out what a page costs and where. Time the paginator and the process's CPU time together:

c0, t0 = time.process_time(), time.perf_counter()
pages = 0
async for page in s3.get_paginator("list_objects_v2").paginate(Bucket="logs"):
    pages += 1
wall, cpu = time.perf_counter() - t0, time.process_time() - c0
print(f"{wall / pages * 1e3:.0f} ms/page wall, {cpu / pages * 1e3:.0f} ms/page client CPU")

Measured: 45 ms per page with no added latency and 93 ms per page with 30 ms added, of which 42–44 ms was CPU time in the client — parsing the 1,000-entry XML response into Python dictionaries, on the event loop. On real S3, page latency is usually higher than here, so the network share will be larger, but the CPU cost per page is the client's own and does not shrink. That CPU cost is what decides how far concurrent listing can go in one process.

Verify: you know the wall time and the client CPU time per page for your bucket.

Listing 100,000 keys with 30 ms per request 5 horizontal bars comparing 1 paginator with the others. Listing 100,000 keys with 30 ms per request 1 paginator 9.07-9.28 s 16 paginators, 1 process 4.26-4.31 s 2 processes 2.22 s 4 processes 1.32 s 8 processes 0.99 s aioboto3 15.5 against SeaweedFS via a 30 ms proxy; 16 prefixes of about 6,250 keys. One process topped out at its CPU budget, about 4.2 s of parsing.

3. List prefixes concurrently

A paginator's pages depend on each other through the continuation token, but different prefixes do not. Discover the top-level prefixes with a delimiter, then list each one with its own paginator:

async def top_prefixes(s3, bucket: str, prefix: str) -> list[str]:
    resp = await s3.list_objects_v2(Bucket=bucket, Prefix=prefix, Delimiter="/")
    return [p["Prefix"] for p in resp.get("CommonPrefixes", [])]

async def count_keys(s3, bucket: str, prefix: str) -> int:
    n = 0
    async for page in s3.get_paginator("list_objects_v2").paginate(Bucket=bucket, Prefix=prefix):
        n += len(page.get("Contents", []))
    return n

prefixes = await top_prefixes(s3, "logs", "events/")          # 16 prefixes in 38 ms
counts = await asyncio.gather(*(count_keys(s3, "logs", p) for p in prefixes))

Measured with 30 ms of latency: 4.26–4.31 s for all 100,000 keys across 16 paginators, against 9.07–9.28 s with one, and 112 pages instead of 100 because each prefix ends with a partial page. The gain stopped at about 2× because the single process was now spending all its time parsing — 4.2 s of CPU in 4.3 s of wall time — and the maximum event-loop lag reached 1.2 s, which in a service means every other request waiting. A layout with no natural prefixes can still be split: list by key ranges with StartAfter, for example one range per leading character.

Verify: concurrent listing returns the same key count as one paginator, and loop lag during it stays within budget.

4. Use processes when parsing is the limit

When listing is CPU-bound in the client, run paginators in several processes, each with its own event loop and client, and combine their results in the parent:

def list_in_process(endpoint: str, prefixes: list[str]) -> int:
    async def run() -> int:
        async with aioboto3.Session().client("s3", endpoint_url=endpoint) as s3:
            counts = await asyncio.gather(*(count_keys(s3, "logs", p) for p in prefixes))
            return sum(counts)
    return asyncio.run(run())

async def count_all(pool: ProcessPoolExecutor, endpoint: str, prefixes: list[str], procs: int) -> int:
    loop = asyncio.get_running_loop()
    groups = [prefixes[i::procs] for i in range(procs)]
    return sum(await asyncio.gather(
        *(loop.run_in_executor(pool, list_in_process, endpoint, g) for g in groups)))

Measured: 2.22 s with two processes, 1.32 s with four and 0.99 s with eight, with the parent's loop lag at 2–6 ms throughout — the parsing happened elsewhere. Return what the parent needs — counts, a filtered subset, or keys written to a file or queue — rather than every key, since sending 100,000 dictionaries back through pickling would cost much of what the processes saved. For bucket-wide inventories that run regularly, S3 Inventory reports produce the full list as files without any listing calls, and are worth considering before building a fast lister.

Verify: the per-process approach returns the same total, and the parent's loop lag stays low while it runs.

Where the time went A grid of 3 rows by 4 columns. Where the time went approach wall client CPU max loop lag 1 paginator 9.07 s 4.37 s (44 ms/page) 75 ms 16 paginators, 1 process 4.31 s 4.21 s 1,221 ms 8 processes, 2 prefixes each 0.99 s in children 6 ms (parent) Botocore parses each 1,000-key XML page in Python on the loop.

5. Resume interrupted listings with StartAfter

A listing of millions of keys can fail part-way — a timeout, a deploy, a crash. Keys come back in UTF-8 binary order, so the last key processed is a complete checkpoint: restart the listing from just after it.

async def list_from(s3, bucket: str, start_after: str | None):
    params = {"Bucket": bucket}
    if start_after:
        params["StartAfter"] = start_after
    async for page in s3.get_paginator("list_objects_v2").paginate(**params):
        contents = page.get("Contents", [])
        for obj in contents:
            yield obj["Key"]
        if contents:
            await save_checkpoint(contents[-1]["Key"])   # after the page's keys are handled

Measured: a listing stopped after 37 pages — 37,000 keys, last key events/5/2025-10-26/0043957.json — and resumed with StartAfter produced 100,000 keys in total, all unique and in sorted order. StartAfter is a key, which stays meaningful indefinitely; a continuation token is opaque and tied to the listing that produced it, so do not store tokens as long-lived checkpoints. Checkpoint after a page's keys have been fully handled, as in checkpointing progress in long-running async jobs, so a crash repeats at most one page.

Verify: an interrupted and resumed listing yields each key exactly once.

How should this bucket be listed? A decision on What limits the listing with 4 outcomes. How should this bucket be listed? What limits the listing? memory stream pages, never collect +3 MiB vs +47 MiB round trips concurrent paginators per prefix 9.1 s to 4.3 s client CPU (loop lag high) paginators in processes 0.99 s with 8 interruptions checkpoint the last key, StartAfter 0 duplicates One process parsed about 23,000 keys a second at most.

Verification

Large listings are handled well when:

  • Pages are streamed, and memory does not grow with key count.
  • Wall time and client CPU per page are measured before parallelising.
  • Prefixes are listed concurrently, and across processes when loop lag shows the client is CPU-bound.
  • Long listings checkpoint the last key and resume with StartAfter.

Diagnostic Hook: when a listing job slows down or a service's latency rises during one, compare the job's CPU time with its wall time. CPU close to wall time means the client is parsing as fast as it can and more concurrency in that process only adds loop lag; CPU far below wall time means the listing is waiting on the network and more prefixes in flight will help.

Pitfalls & edge cases

  • Collecting every key before processing. Measured: 47 MiB per 100,000 keys.
  • Assuming listing is network-bound. Botocore parsing took 42–44 ms of CPU per page.
  • Many paginators on a service's loop. Measured: 1.2 s of loop lag.
  • Storing continuation tokens as checkpoints. Use the last key and StartAfter.

Frequently Asked Questions

How do I list S3 objects with aioboto3?

Use client.get_paginator("list_objects_v2") and iterate its pages with async for, handling each page's Contents as it arrives; streaming 100,000 keys added 3 MiB of memory in testing.

How can I list a large S3 bucket faster?

List prefixes concurrently: 16 paginators took 4.3 s instead of 9.1 s for 100,000 keys. Beyond that, botocore's XML parsing (about 42 ms of CPU per page) limits one process, and 8 processes took 0.99 s.

Why does listing S3 objects block my event loop?

Each 1,000-key response is parsed in Python on the loop. Sixteen concurrent paginators caused 1.2 s of loop lag; run large listings in separate processes.

How do I resume an interrupted S3 listing?

Save the last key you processed and pass it as StartAfter; a resumed listing returned all 100,000 keys exactly once.