Showing Progress for Concurrent Tasks with rich and tqdm¶
A script that runs two hundred downloads concurrently needs to show that it is alive. Progress bars are the obvious answer, and both popular libraries — rich and tqdm — work from asyncio code, because their update calls are plain, fast function calls. The cost hides in one place. Measured on Python 3.14 with rich 15.0.0 and tqdm 4.70.1, 200 tasks each advancing 50 times: with no progress display the run took 0.136 s; with one rich bar for the whole job 0.144 s; with one tqdm bar 0.135 s. Giving each task its own rich bar, added as the task started, took 2.66 s — and the event loop stalled for 2.38 s in one stretch, because Progress.add_task re-renders every bar synchronously on the calling thread, which here is the event loop. Adding the same 200 bars before starting the display took 0.20 s. A single update() call cost 0.50 µs in rich and 0.12 µs in tqdm. This guide shows progress without paying for it.
Prerequisites¶
- Python 3.11+,
pip install richand/orpip install tqdm. - Concurrent fan-out, from downloading many URLs concurrently with progress.
- Loop lag as a health signal, from measuring event loop lag in production.
1. Use one bar for the whole job¶
The cheapest progress display counts finished units across all tasks. Updates are synchronous calls with no await, so they cannot interleave badly:
import asyncio
from rich.progress import BarColumn, Progress, TaskProgressColumn, TextColumn, TimeRemainingColumn
async def worker(url: str, advance) -> None:
for chunk in await fetch_chunks(url): # your real work
await process(chunk)
advance()
async def main(urls: list[str]) -> None:
with Progress(TextColumn("{task.description}"), BarColumn(), TaskProgressColumn(),
TimeRemainingColumn(), refresh_per_second=10) as progress:
overall = progress.add_task("downloading", total=len(urls) * CHUNKS)
await asyncio.gather(*(worker(u, lambda: progress.update(overall, advance=1))
for u in urls))
Measured: 200 tasks × 50 updates took 0.144 s with the bar against 0.136 s without it, and the loop-lag probe's p99 stayed between 0.86 and 1.25 ms in both cases. rich redraws from its own thread ten times a second; update() only changes counters under a lock. refresh_per_second bounds the redraw cost no matter how often tasks call update().
Verify: wall time and loop lag with the progress bar are within a few percent of a run without it.
2. Avoid adding many bars to a live display¶
Per-task bars feel natural — one line per download — and are the expensive case. Timing each call showed where the time went:
with Progress(refresh_per_second=10) as p: # display is live
async def one(i: int) -> None:
tid = p.add_task(f"item {i}", total=50) # re-renders every bar, on the loop
for _ in range(50):
await asyncio.sleep(0.002)
p.update(tid, advance=1) # ~0.5 us
await asyncio.gather(*(one(i) for i in range(200)))
Measured: the 200 add_task calls took 2.46 s in total, up to 45.9 ms each, while 10,000 update calls took 0.010 s altogether. add_task calls refresh() when the display is running, and a refresh renders every existing row, so the cost grows with the number of bars already present — quadratic over the run. All of it happens on the thread that called add_task, which in an async program is the event loop: every other task, timer and socket waits. Two fixes, both measured: create the bars before progress.start() (0.20 s for the same run), or keep the visible set small.
Verify: time the slowest add_task call in your script; if it exceeds a few milliseconds, the display has too many rows.
3. Show an overall bar plus a few active rows¶
For long-running items, users want to see what is in flight. Bound the visible rows to the concurrency limit and remove each row when its item completes:
async def main(items: list[str], concurrency: int = 8) -> None:
slots = asyncio.Semaphore(concurrency)
with Progress(refresh_per_second=10) as p:
overall = p.add_task("all items", total=len(items))
async def one(item: str) -> None:
async with slots:
row = p.add_task(item, total=STEPS)
try:
for _ in range(STEPS):
await step(item)
p.update(row, advance=1)
finally:
p.remove_task(row)
p.update(overall, advance=1)
await asyncio.gather(*(one(i) for i in items))
Measured with eight concurrent rows plus the overall bar: the slowest add_task took 2.15 ms and all 200 together 253 ms, over a run of 2.82 s against 2.62 s without per-item rows. That overhead is acceptable for items that take seconds each; for millisecond items, drop the per-item rows and keep only the overall bar. The try/finally keeps a cancelled or failed item from leaving a frozen row on screen.
Verify: the number of visible rows never exceeds the concurrency limit plus one.
4. Use tqdm's asyncio helpers for simple fan-outs¶
tqdm ships drop-in replacements for asyncio.gather and asyncio.as_completed that count completed awaitables:
from tqdm.asyncio import tqdm
async def main() -> None:
results = await tqdm.gather(*(job(i) for i in range(2000)), desc="jobs")
# results are in submission order, like asyncio.gather
for fut in tqdm.as_completed([job(i) for i in range(2000)], desc="jobs"):
result = await fut # completion order
Measured over 2,000 short jobs: tqdm.gather took 0.030 s against 0.025 s for asyncio.gather, and returned results in submission order; tqdm.as_completed yielded results in completion order (the first three were jobs 1380, 950 and 1390). tqdm's per-update cost was the lowest measured, 0.12 µs. It has no equivalent of rich's multi-row layouts, which makes it the better fit for "one progress line per script".
Verify: the bar reaches 100% exactly when the gather returns.
5. Keep progress output out of logs and pipes¶
Progress bars are for terminals. In CI logs, cron mail and pipes they become thousands of carriage-return-separated lines. Both libraries can detect this:
import sys
interactive = sys.stderr.isatty()
with Progress(disable=not interactive) as p: # rich: no output when not a TTY
...
for x in tqdm(items, disable=None): # tqdm: None means "disable if not a TTY"
...
When the bar is disabled, log a summary line at intervals instead — completed count and rate every few seconds — from a periodic task, as described in running periodic tasks without drift. A side effect worth knowing: while a rich Progress is live, print() output is redirected through rich's console so it appears above the bars; in the measurements here, prints from inside the display went into the test console rather than the terminal. Use progress.console.print() for messages that should appear alongside the bars.
Verify: running the script with 2>&1 | cat produces readable log lines, not a stream of redraws.
Verification¶
Progress display is safe in an async script when:
- Loop lag and run time with the display are close to a run without it.
- The number of rich rows added while live is bounded by the concurrency limit, or rows are created before
start(). - Rows are removed in
finally, so failures and cancellations do not leave stale rows. - Bars are disabled when output is not a terminal, with periodic summary logging instead.
Diagnostic Hook: run a loop-lag probe (an asyncio.sleep(0.005) loop that records overshoot) alongside the progress display during development. A single multi-second lag sample near the start of the run is the signature of many add_task calls on a live display.
Pitfalls & edge cases¶
- Per-task rich bars on a live display. Measured: 2.66 s instead of 0.14 s, with a 2.38 s stall.
- Forgetting
remove_task. Finished and failed rows accumulate and slow every later refresh. - Bars in CI logs. Use
disable=with a TTY check. print()while rich is live. It is redirected through rich's console.
Frequently Asked Questions¶
How do I show a progress bar for asyncio.gather?
Use one rich Progress task or tqdm bar for the whole job and call update(advance=1) as each unit finishes, or use tqdm.asyncio's tqdm.gather as a drop-in replacement. Measured overhead was about 0.01 s on a 0.14 s run of 10,000 updates.
Why does my rich progress bar make asyncio slow?
Adding bars to a running display calls refresh() synchronously, rendering every row on the event loop thread. With 200 bars added while live, add_task took up to 45.9 ms per call and stalled the loop for 2.38 s. Add bars before start(), or keep only a few visible.
Is tqdm or rich faster for async progress?
tqdm's update() cost 0.12 µs and rich's 0.50 µs; both were negligible next to real work. rich is better for multi-row layouts; tqdm has asyncio drop-ins for gather and as_completed.
Are progress bar updates thread-safe in asyncio?
Updates from coroutines all run on the event loop thread, so there is no race; rich additionally guards its state with a lock because it redraws from a background thread.
Related¶
- Async Scripts & CLIs — up to the topic overview.
- Reading stdin asynchronously — the input side of a pipeline script.
- Asyncio Fundamentals & Event Loop Architecture — the section overview.