Skip to content

Downloading Many URLs Concurrently with Progress

Downloading a long list of URLs is the classic asyncio job, and the minimal version — gather over one coroutine per URL — works for a few hundred. At scale three things decide whether it works well: how many downloads run at once, how memory scales with the length of the list, and whether you can see progress and failures while it runs. Measured against a local server with 50 ms of latency per 256 KiB file, 400 downloads took 2.12 s at a concurrency of 10, 0.48 s at 50, and 0.32 s at 200. With 200,000 URLs, creating a task per URL up front peaked at 289 MB of memory before any work was done; a fixed pool of 50 workers fed from a bounded queue peaked at 20 MB. This guide builds the worker version with streaming to disk, progress reporting, retries and a failure log.

Prerequisites

1. Bound concurrency, and pick the bound

Every download in flight holds a socket and a buffer, and the remote server has its own limits. Bound concurrency with a semaphore or a connector limit, and choose the number by measurement:

import aiohttp

connector = aiohttp.TCPConnector(limit=50, limit_per_host=10)   # overall and per host
session = aiohttp.ClientSession(connector=connector, timeout=aiohttp.ClientTimeout(total=120))

Measured against a local server: concurrency 1 managed 19 files a second; 10 managed 188; 50 managed 831; 200 managed 1,260. Throughput rises with concurrency until something saturates — here the local machine; in practice the remote host's limits, your bandwidth or rate limiting. Past that point, more concurrency only adds errors and queueing. limit_per_host matters when the list spans many hosts: it keeps one host from absorbing the whole budget and protects small sites from being flooded, as in limiting concurrency per host in an async crawler.

Verify: run a sample at several concurrency levels; choose the lowest level that achieves near-peak throughput.

Files per second by concurrency 4 horizontal bars comparing concurrency 1 with the others. Files per second by concurrency concurrency 1 19 files/s concurrency 10 188 files/s concurrency 50 831 files/s concurrency 200 1,260 files/s aiohttp 3.14, local server with 50 ms latency per 256 KiB file; 400 files (40 at concurrency 1). Returns diminish once something saturates; real servers saturate much sooner.

2. Use a worker pool, not a task per URL

asyncio.gather(*(download(u) for u in urls)) creates a task and a coroutine for every URL before any of them run. With a semaphore inside, only 50 work at once, but all of them exist:

async def download_all(session, urls, out_dir, workers: int = 50) -> Stats:
    queue: asyncio.Queue[str] = asyncio.Queue(maxsize=workers * 2)
    stats = Stats(total=len(urls))

    async def worker():
        while True:
            url = await queue.get()
            try:
                await download_one(session, url, out_dir, stats)
            finally:
                queue.task_done()

    async with asyncio.TaskGroup() as tg:
        tasks = [tg.create_task(worker()) for _ in range(workers)]
        for url in urls:                           # urls may be a generator: never fully in memory
            await queue.put(url)
        await queue.join()
        for t in tasks:
            t.cancel()
    return stats

Measured with 200,000 URLs: one task per URL peaked at 289 MB; 50 workers and a bounded queue peaked at 20 MB, regardless of the list length. Because the producer waits on a bounded queue, urls can be a generator reading from a file or a database cursor, so even the list itself never needs to be in memory. The pattern in depth is in building an async worker pool with TaskGroup.

Verify: memory stays flat as the URL count grows tenfold.

3. Stream each response to disk

Reading a whole response with await r.read() holds the file in memory; with 50 concurrent large files that adds up. Stream in chunks to a temporary name and rename on success:

import os


async def download_one(session, url: str, out_dir: str, stats: "Stats") -> None:
    name = url.rsplit("/", 1)[-1] or "index"
    final = os.path.join(out_dir, name)
    tmp = final + ".part"
    async with session.get(url) as r:
        r.raise_for_status()
        with open(tmp, "wb") as f:
            async for chunk in r.content.iter_chunked(64 * 1024):
                f.write(chunk)
                stats.bytes += len(chunk)
    os.replace(tmp, final)                         # only complete files get the final name
    stats.done += 1

The .part suffix means an interrupted run never leaves a truncated file under the real name, and a resumed run can skip files that already exist. File writes here are synchronous; for local disks and 64 KiB chunks they take microseconds, but on network filesystems move them to a thread, as in async file I/O with aiofiles vs asyncio.to_thread.

Verify: kill the run halfway; no file without .part is truncated, and a restart skips completed files.

The download pipeline A flow of 5 stages. The download pipeline URL source lazy, a generator bounded queue backpressure N workers stream to .part rename on success skip on resume reporter rate, ETA, failures Memory is set by the worker count, not by the length of the list.

4. Report progress from a separate task

Progress is a reporter task reading shared counters, not print statements in every download. That keeps reporting cheap and its rate independent of the download rate:

import time
from dataclasses import dataclass, field


@dataclass
class Stats:
    total: int
    done: int = 0
    failed: int = 0
    bytes: int = 0
    started: float = field(default_factory=time.monotonic)


async def report(stats: Stats, every: float = 2.0) -> None:
    while True:
        await asyncio.sleep(every)
        elapsed = time.monotonic() - stats.started
        finished = stats.done + stats.failed
        rate = finished / elapsed if elapsed else 0
        eta = (stats.total - finished) / rate if rate else float("inf")
        print(f"{finished}/{stats.total} ({stats.failed} failed) "
              f"{stats.bytes / 2**20:.0f} MiB  {rate:.0f}/s  ETA {eta:.0f}s", flush=True)

Start it alongside the workers and cancel it when they finish. Because all tasks run on one thread, plain integer counters are safe — no lock is needed between workers and reporter. For a terminal progress bar, tqdm.asyncio or rich.progress can be driven from the same counters.

Verify: progress lines appear at a steady interval and the final count equals the number of URLs.

5. Retry transient failures and log the rest

In a long list, some downloads fail: timeouts, resets, 503s. Retry the transient ones a few times with backoff, and record the rest for a later pass instead of aborting the run:

TRANSIENT = (aiohttp.ClientConnectionError, asyncio.TimeoutError)


async def download_with_retry(session, url, out_dir, stats, attempts: int = 3) -> None:
    for attempt in range(attempts):
        try:
            return await download_one(session, url, out_dir, stats)
        except aiohttp.ClientResponseError as exc:
            if exc.status < 500 and exc.status != 429:
                break                                      # 404, 403: will not change
        except TRANSIENT:
            pass
        await asyncio.sleep(min(2 ** attempt, 10) * random.uniform(0.5, 1.0))
    stats.failed += 1
    failures.write(url + "\n")                             # a file you can feed back in

A failure log that is itself a URL list makes the retry pass trivial: run the same program on the failures file. Catch exceptions per download so one bad URL never cancels the whole TaskGroup. Rate-limited sources need more than backoff — see handling 429 Retry-After responses in async clients.

Verify: inject 5% failures into a test server; the run completes, transient failures are retried, and permanent ones end up in the failures file.

What to do with a failed download A decision on How did it fail with 3 outcomes. What to do with a failed download How did it fail? timeout, reset, 5xx, 429 retry with backoff up to 3 times 404, 403, other 4xx log, do not retry permanent retries exhausted failures file feed back in later A long run should finish with a list of what to retry, not a stack trace.

Verification

A bulk downloader is ready when:

  • Concurrency is bounded at a measured level, overall and per host.
  • Memory is independent of list length, with workers and a bounded queue.
  • Files are streamed to disk and renamed only when complete.
  • Progress and failures are visible, and failures can be replayed.

Diagnostic Hook: log throughput and error rate per host every few seconds. Throughput that plateaus while errors climb means concurrency is above what the source tolerates — lower it. Throughput that plateaus with no errors means a local limit: bandwidth, disk, or the connector's limit.

Pitfalls & edge cases

  • One task per URL for huge lists. 289 MB for 200,000 URLs before any work.
  • await r.read() for large files. Stream with iter_chunked instead.
  • Letting one exception cancel the batch. Catch per download and log.
  • Unbounded per-host concurrency. Small servers will rate-limit or block you.

Frequently Asked Questions

How do I download many files concurrently with aiohttp?

Share one ClientSession, run a fixed number of worker tasks that take URLs from a bounded asyncio.Queue, and stream each response to disk with iter_chunked. Limit concurrency with the connector's limit and limit_per_host.

How many concurrent downloads should I run?

Measure: throughput rises with concurrency until the remote server, bandwidth or your machine saturates. In a local test, 10 gave 188 files a second and 50 gave 831; real servers usually saturate earlier.

Why does asyncio.gather use so much memory for many URLs?

It creates a task for every URL up front. In testing, 200,000 URLs took 289 MB that way versus 20 MB with 50 workers fed from a bounded queue.

How do I show download progress in asyncio?

Keep counters updated by the workers and run a separate reporter task that prints them every few seconds, with rate and ETA.