Skip to content

Showing Progress for Concurrent Tasks with rich and tqdm

A script that runs two hundred downloads concurrently needs to show that it is alive. Progress bars are the obvious answer, and both popular libraries — rich and tqdm — work from asyncio code, because their update calls are plain, fast function calls. The cost hides in one place. Measured on Python 3.14 with rich 15.0.0 and tqdm 4.70.1, 200 tasks each advancing 50 times: with no progress display the run took 0.136 s; with one rich bar for the whole job 0.144 s; with one tqdm bar 0.135 s. Giving each task its own rich bar, added as the task started, took 2.66 s — and the event loop stalled for 2.38 s in one stretch, because Progress.add_task re-renders every bar synchronously on the calling thread, which here is the event loop. Adding the same 200 bars before starting the display took 0.20 s. A single update() call cost 0.50 µs in rich and 0.12 µs in tqdm. This guide shows progress without paying for it.

Prerequisites

1. Use one bar for the whole job

The cheapest progress display counts finished units across all tasks. Updates are synchronous calls with no await, so they cannot interleave badly:

import asyncio
from rich.progress import BarColumn, Progress, TaskProgressColumn, TextColumn, TimeRemainingColumn


async def worker(url: str, advance) -> None:
    for chunk in await fetch_chunks(url):          # your real work
        await process(chunk)
        advance()


async def main(urls: list[str]) -> None:
    with Progress(TextColumn("{task.description}"), BarColumn(), TaskProgressColumn(),
                  TimeRemainingColumn(), refresh_per_second=10) as progress:
        overall = progress.add_task("downloading", total=len(urls) * CHUNKS)
        await asyncio.gather(*(worker(u, lambda: progress.update(overall, advance=1))
                               for u in urls))

Measured: 200 tasks × 50 updates took 0.144 s with the bar against 0.136 s without it, and the loop-lag probe's p99 stayed between 0.86 and 1.25 ms in both cases. rich redraws from its own thread ten times a second; update() only changes counters under a lock. refresh_per_second bounds the redraw cost no matter how often tasks call update().

Verify: wall time and loop lag with the progress bar are within a few percent of a run without it.

200 tasks x 50 updates, total run time 5 horizontal bars comparing no progress display with the others. 200 tasks x 50 updates, total run time no progress display 0.136 s one tqdm bar 0.135 s one rich bar 0.144 s 200 rich bars, added before start 0.20 s 200 rich bars, added while running 2.66 s rich 15.0.0, tqdm 4.70.1, Python 3.14; each step awaited a 2 ms sleep. Progress itself is cheap; adding rich bars to a live display is not.

2. Avoid adding many bars to a live display

Per-task bars feel natural — one line per download — and are the expensive case. Timing each call showed where the time went:

with Progress(refresh_per_second=10) as p:            # display is live
    async def one(i: int) -> None:
        tid = p.add_task(f"item {i}", total=50)       # re-renders every bar, on the loop
        for _ in range(50):
            await asyncio.sleep(0.002)
            p.update(tid, advance=1)                  # ~0.5 us
    await asyncio.gather(*(one(i) for i in range(200)))

Measured: the 200 add_task calls took 2.46 s in total, up to 45.9 ms each, while 10,000 update calls took 0.010 s altogether. add_task calls refresh() when the display is running, and a refresh renders every existing row, so the cost grows with the number of bars already present — quadratic over the run. All of it happens on the thread that called add_task, which in an async program is the event loop: every other task, timer and socket waits. Two fixes, both measured: create the bars before progress.start() (0.20 s for the same run), or keep the visible set small.

Verify: time the slowest add_task call in your script; if it exceeds a few milliseconds, the display has too many rows.

3. Show an overall bar plus a few active rows

For long-running items, users want to see what is in flight. Bound the visible rows to the concurrency limit and remove each row when its item completes:

async def main(items: list[str], concurrency: int = 8) -> None:
    slots = asyncio.Semaphore(concurrency)
    with Progress(refresh_per_second=10) as p:
        overall = p.add_task("all items", total=len(items))

        async def one(item: str) -> None:
            async with slots:
                row = p.add_task(item, total=STEPS)
                try:
                    for _ in range(STEPS):
                        await step(item)
                        p.update(row, advance=1)
                finally:
                    p.remove_task(row)
            p.update(overall, advance=1)

        await asyncio.gather(*(one(i) for i in items))

Measured with eight concurrent rows plus the overall bar: the slowest add_task took 2.15 ms and all 200 together 253 ms, over a run of 2.82 s against 2.62 s without per-item rows. That overhead is acceptable for items that take seconds each; for millisecond items, drop the per-item rows and keep only the overall bar. The try/finally keeps a cancelled or failed item from leaving a frozen row on screen.

Verify: the number of visible rows never exceeds the concurrency limit plus one.

Where rich's time went, 200 items A grid of 3 rows by 4 columns. Where rich's time went, 200 items pattern add_task total slowest add_task update() total 200 bars added while live 2.46 s 45.9 ms 0.010 s overall + at most 8 item rows 253 ms 2.15 ms small one overall bar one call negligible ~0.5 us per call add_task re-renders all rows; its cost scales with how many are visible.

4. Use tqdm's asyncio helpers for simple fan-outs

tqdm ships drop-in replacements for asyncio.gather and asyncio.as_completed that count completed awaitables:

from tqdm.asyncio import tqdm


async def main() -> None:
    results = await tqdm.gather(*(job(i) for i in range(2000)), desc="jobs")
    # results are in submission order, like asyncio.gather

    for fut in tqdm.as_completed([job(i) for i in range(2000)], desc="jobs"):
        result = await fut                       # completion order

Measured over 2,000 short jobs: tqdm.gather took 0.030 s against 0.025 s for asyncio.gather, and returned results in submission order; tqdm.as_completed yielded results in completion order (the first three were jobs 1380, 950 and 1390). tqdm's per-update cost was the lowest measured, 0.12 µs. It has no equivalent of rich's multi-row layouts, which makes it the better fit for "one progress line per script".

Verify: the bar reaches 100% exactly when the gather returns.

5. Keep progress output out of logs and pipes

Progress bars are for terminals. In CI logs, cron mail and pipes they become thousands of carriage-return-separated lines. Both libraries can detect this:

import sys

interactive = sys.stderr.isatty()

with Progress(disable=not interactive) as p:      # rich: no output when not a TTY
    ...

for x in tqdm(items, disable=None):               # tqdm: None means "disable if not a TTY"
    ...

When the bar is disabled, log a summary line at intervals instead — completed count and rate every few seconds — from a periodic task, as described in running periodic tasks without drift. A side effect worth knowing: while a rich Progress is live, print() output is redirected through rich's console so it appears above the bars; in the measurements here, prints from inside the display went into the test console rather than the terminal. Use progress.console.print() for messages that should appear alongside the bars.

Verify: running the script with 2>&1 | cat produces readable log lines, not a stream of redraws.

What kind of progress display fits this script? A decision on What is being tracked with 4 outcomes. What kind of progress display fits this script? What is being tracked? many short items one overall bar cost: microseconds a few long items at a time overall + rows for active items rows <= concurrency a plain gather or as_completed tqdm.gather / tqdm.as_completed drop-in not a terminal disable bars, log summaries isatty() check Never add hundreds of rich rows to a live display from the event loop.

Verification

Progress display is safe in an async script when:

  • Loop lag and run time with the display are close to a run without it.
  • The number of rich rows added while live is bounded by the concurrency limit, or rows are created before start().
  • Rows are removed in finally, so failures and cancellations do not leave stale rows.
  • Bars are disabled when output is not a terminal, with periodic summary logging instead.

Diagnostic Hook: run a loop-lag probe (an asyncio.sleep(0.005) loop that records overshoot) alongside the progress display during development. A single multi-second lag sample near the start of the run is the signature of many add_task calls on a live display.

Pitfalls & edge cases

  • Per-task rich bars on a live display. Measured: 2.66 s instead of 0.14 s, with a 2.38 s stall.
  • Forgetting remove_task. Finished and failed rows accumulate and slow every later refresh.
  • Bars in CI logs. Use disable= with a TTY check.
  • print() while rich is live. It is redirected through rich's console.

Frequently Asked Questions

How do I show a progress bar for asyncio.gather?

Use one rich Progress task or tqdm bar for the whole job and call update(advance=1) as each unit finishes, or use tqdm.asyncio's tqdm.gather as a drop-in replacement. Measured overhead was about 0.01 s on a 0.14 s run of 10,000 updates.

Why does my rich progress bar make asyncio slow?

Adding bars to a running display calls refresh() synchronously, rendering every row on the event loop thread. With 200 bars added while live, add_task took up to 45.9 ms per call and stalled the loop for 2.38 s. Add bars before start(), or keep only a few visible.

Is tqdm or rich faster for async progress?

tqdm's update() cost 0.12 µs and rich's 0.50 µs; both were negligible next to real work. rich is better for multi-row layouts; tqdm has asyncio drop-ins for gather and as_completed.

Are progress bar updates thread-safe in asyncio?

Updates from coroutines all run on the event loop thread, so there is no race; rich additionally guards its state with a lock because it redraws from a background thread.