Downloading Many URLs Concurrently with Progress¶
Downloading a long list of URLs is the classic asyncio job, and the minimal version — gather over one coroutine per URL — works for a few hundred. At scale three things decide whether it works well: how many downloads run at once, how memory scales with the length of the list, and whether you can see progress and failures while it runs. Measured against a local server with 50 ms of latency per 256 KiB file, 400 downloads took 2.12 s at a concurrency of 10, 0.48 s at 50, and 0.32 s at 200. With 200,000 URLs, creating a task per URL up front peaked at 289 MB of memory before any work was done; a fixed pool of 50 workers fed from a bounded queue peaked at 20 MB. This guide builds the worker version with streaming to disk, progress reporting, retries and a failure log.
Prerequisites¶
- Python 3.11+,
pip install aiohttp; measured with aiohttp 3.14. - Bounded concurrency, from limiting concurrent requests with asyncio.Semaphore.
- Session reuse, from reusing aiohttp ClientSession across requests.
1. Bound concurrency, and pick the bound¶
Every download in flight holds a socket and a buffer, and the remote server has its own limits. Bound concurrency with a semaphore or a connector limit, and choose the number by measurement:
import aiohttp
connector = aiohttp.TCPConnector(limit=50, limit_per_host=10) # overall and per host
session = aiohttp.ClientSession(connector=connector, timeout=aiohttp.ClientTimeout(total=120))
Measured against a local server: concurrency 1 managed 19 files a second; 10 managed 188; 50 managed 831; 200 managed 1,260. Throughput rises with concurrency until something saturates — here the local machine; in practice the remote host's limits, your bandwidth or rate limiting. Past that point, more concurrency only adds errors and queueing. limit_per_host matters when the list spans many hosts: it keeps one host from absorbing the whole budget and protects small sites from being flooded, as in limiting concurrency per host in an async crawler.
Verify: run a sample at several concurrency levels; choose the lowest level that achieves near-peak throughput.
2. Use a worker pool, not a task per URL¶
asyncio.gather(*(download(u) for u in urls)) creates a task and a coroutine for every URL before any of them run. With a semaphore inside, only 50 work at once, but all of them exist:
async def download_all(session, urls, out_dir, workers: int = 50) -> Stats:
queue: asyncio.Queue[str] = asyncio.Queue(maxsize=workers * 2)
stats = Stats(total=len(urls))
async def worker():
while True:
url = await queue.get()
try:
await download_one(session, url, out_dir, stats)
finally:
queue.task_done()
async with asyncio.TaskGroup() as tg:
tasks = [tg.create_task(worker()) for _ in range(workers)]
for url in urls: # urls may be a generator: never fully in memory
await queue.put(url)
await queue.join()
for t in tasks:
t.cancel()
return stats
Measured with 200,000 URLs: one task per URL peaked at 289 MB; 50 workers and a bounded queue peaked at 20 MB, regardless of the list length. Because the producer waits on a bounded queue, urls can be a generator reading from a file or a database cursor, so even the list itself never needs to be in memory. The pattern in depth is in building an async worker pool with TaskGroup.
Verify: memory stays flat as the URL count grows tenfold.
3. Stream each response to disk¶
Reading a whole response with await r.read() holds the file in memory; with 50 concurrent large files that adds up. Stream in chunks to a temporary name and rename on success:
import os
async def download_one(session, url: str, out_dir: str, stats: "Stats") -> None:
name = url.rsplit("/", 1)[-1] or "index"
final = os.path.join(out_dir, name)
tmp = final + ".part"
async with session.get(url) as r:
r.raise_for_status()
with open(tmp, "wb") as f:
async for chunk in r.content.iter_chunked(64 * 1024):
f.write(chunk)
stats.bytes += len(chunk)
os.replace(tmp, final) # only complete files get the final name
stats.done += 1
The .part suffix means an interrupted run never leaves a truncated file under the real name, and a resumed run can skip files that already exist. File writes here are synchronous; for local disks and 64 KiB chunks they take microseconds, but on network filesystems move them to a thread, as in async file I/O with aiofiles vs asyncio.to_thread.
Verify: kill the run halfway; no file without .part is truncated, and a restart skips completed files.
4. Report progress from a separate task¶
Progress is a reporter task reading shared counters, not print statements in every download. That keeps reporting cheap and its rate independent of the download rate:
import time
from dataclasses import dataclass, field
@dataclass
class Stats:
total: int
done: int = 0
failed: int = 0
bytes: int = 0
started: float = field(default_factory=time.monotonic)
async def report(stats: Stats, every: float = 2.0) -> None:
while True:
await asyncio.sleep(every)
elapsed = time.monotonic() - stats.started
finished = stats.done + stats.failed
rate = finished / elapsed if elapsed else 0
eta = (stats.total - finished) / rate if rate else float("inf")
print(f"{finished}/{stats.total} ({stats.failed} failed) "
f"{stats.bytes / 2**20:.0f} MiB {rate:.0f}/s ETA {eta:.0f}s", flush=True)
Start it alongside the workers and cancel it when they finish. Because all tasks run on one thread, plain integer counters are safe — no lock is needed between workers and reporter. For a terminal progress bar, tqdm.asyncio or rich.progress can be driven from the same counters.
Verify: progress lines appear at a steady interval and the final count equals the number of URLs.
5. Retry transient failures and log the rest¶
In a long list, some downloads fail: timeouts, resets, 503s. Retry the transient ones a few times with backoff, and record the rest for a later pass instead of aborting the run:
TRANSIENT = (aiohttp.ClientConnectionError, asyncio.TimeoutError)
async def download_with_retry(session, url, out_dir, stats, attempts: int = 3) -> None:
for attempt in range(attempts):
try:
return await download_one(session, url, out_dir, stats)
except aiohttp.ClientResponseError as exc:
if exc.status < 500 and exc.status != 429:
break # 404, 403: will not change
except TRANSIENT:
pass
await asyncio.sleep(min(2 ** attempt, 10) * random.uniform(0.5, 1.0))
stats.failed += 1
failures.write(url + "\n") # a file you can feed back in
A failure log that is itself a URL list makes the retry pass trivial: run the same program on the failures file. Catch exceptions per download so one bad URL never cancels the whole TaskGroup. Rate-limited sources need more than backoff — see handling 429 Retry-After responses in async clients.
Verify: inject 5% failures into a test server; the run completes, transient failures are retried, and permanent ones end up in the failures file.
Verification¶
A bulk downloader is ready when:
- Concurrency is bounded at a measured level, overall and per host.
- Memory is independent of list length, with workers and a bounded queue.
- Files are streamed to disk and renamed only when complete.
- Progress and failures are visible, and failures can be replayed.
Diagnostic Hook: log throughput and error rate per host every few seconds. Throughput that plateaus while errors climb means concurrency is above what the source tolerates — lower it. Throughput that plateaus with no errors means a local limit: bandwidth, disk, or the connector's limit.
Pitfalls & edge cases¶
- One task per URL for huge lists. 289 MB for 200,000 URLs before any work.
await r.read()for large files. Stream withiter_chunkedinstead.- Letting one exception cancel the batch. Catch per download and log.
- Unbounded per-host concurrency. Small servers will rate-limit or block you.
Frequently Asked Questions¶
How do I download many files concurrently with aiohttp?
Share one ClientSession, run a fixed number of worker tasks that take URLs from a bounded asyncio.Queue, and stream each response to disk with iter_chunked. Limit concurrency with the connector's limit and limit_per_host.
How many concurrent downloads should I run?
Measure: throughput rises with concurrency until the remote server, bandwidth or your machine saturates. In a local test, 10 gave 188 files a second and 50 gave 831; real servers usually saturate earlier.
Why does asyncio.gather use so much memory for many URLs?
It creates a task for every URL up front. In testing, 200,000 URLs took 289 MB that way versus 20 MB with 50 workers fed from a bounded queue.
How do I show download progress in asyncio?
Keep counters updated by the workers and run a separate reporter task that prints them every few seconds, with rate and ETA.
Related¶
- Async HTTP Clients & Servers — up to the topic overview.
- Downloading many S3 objects concurrently — the same pattern against object storage.
- Network I/O & Protocol Handling — the section overview.