Choosing Concurrency for a Command-Line Tool¶
A command-line tool has different priorities from a server: it starts on every invocation, runs once, and the person running it presses Ctrl-C when it takes too long. Those three facts change which concurrency model fits. Measured on Python 3.14 with a tool fetching URLs from a local server that answers in 50 ms: fetching 200 URLs took 10.25 s one at a time, 0.39 s with 32 threads and 0.37 s with asyncio and 32 concurrent requests — but importing aiohttp alone added 157 ms to start-up, against 34 ms for urllib.request. With 2,000 URLs and 200 concurrent requests, asyncio took 1.09 s against 0.63 s for threads until aiohttp's default cap of 100 connections was raised. When the user pressed Ctrl-C during slow requests, the asyncio version exited in 0.05 s; the threaded version exited only after 4.12 s, when the requests already in flight finished. For CPU work, 8 threads gave a 4.8× speed-up hashing files but none on pure-Python processing, where a process pool gave 4.7×. This guide measures each factor and turns them into a choice.
Prerequisites¶
- Python 3.11+; aiohttp or httpx if comparing async HTTP.
- Signal handling, from handling Ctrl-C in asyncio scripts.
- The topic overview, Threading vs Multiprocessing vs asyncio.
1. Measure start-up cost¶
A server pays its import time once; a CLI pays it on every run. Time a fresh interpreter importing each candidate, repeated to get a stable median:
for stmt in ["pass", "import asyncio", "import concurrent.futures",
"import urllib.request", "import requests", "import httpx", "import aiohttp"]:
times = []
for _ in range(15):
t = time.perf_counter()
subprocess.run([sys.executable, "-c", stmt], check=True)
times.append(time.perf_counter() - t)
print(stmt, statistics.median(times))
Measured, including interpreter start: an empty program 9.0 ms; concurrent.futures 26.1 ms; urllib.request 34.2 ms; asyncio 42.4 ms; requests 81.7 ms; httpx 89.3 ms; aiohttp 156.5 ms. For a tool that makes three requests and exits, aiohttp's import costs more than the requests. For a tool that fetches thousands, it does not matter. python -X importtime breaks the number down by module, and deferring a heavy import into the function that needs it keeps --help fast.
Verify: time mytool --help stays under about 100 ms, and heavy imports happen only on code paths that use them.
2. Compare throughput at the tool's real size¶
Write the I/O loop both ways. Threads with the standard library:
def get(i):
with urllib.request.urlopen(f"{BASE}/item/{i}", timeout=30) as r:
return len(r.read())
with ThreadPoolExecutor(max_workers=32) as ex:
total = sum(ex.map(get, range(n)))
And asyncio with aiohttp:
async def main(n, limit):
sem = asyncio.Semaphore(limit)
connector = aiohttp.TCPConnector(limit=limit) # default limit is 100
async with aiohttp.ClientSession(connector=connector) as s:
async def get(i):
async with sem, s.get(f"{BASE}/item/{i}") as r:
return len(await r.read())
return sum(await asyncio.gather(*(get(i) for i in range(n))))
Measured with 200 URLs: sequential 10.25 s; 32 threads 0.39 s; asyncio with 32 at a time 0.37 s. With 2,000 URLs at 32 at a time, both took about 3.3 s. With 200 at a time, threads took 0.63 s, while asyncio with a Semaphore(200) and a default ClientSession took 1.09 s — aiohttp's TCPConnector allows 100 connections by default, which halved the effective concurrency. Setting limit=200 on the connector brought asyncio to 0.62 s. At these sizes, the two models delivered the same throughput; the difference was a library default.
Verify: run the tool against its realistic input size and concurrency, and check that every layer's limit — semaphore, connector, pool — matches the intended concurrency.
3. Test what Ctrl-C does¶
A server is stopped by its supervisor; a CLI is stopped by a person who wants it to stop now. Send SIGINT one second into a run where each request takes 5 seconds:
p = subprocess.Popen([sys.executable, "tool.py"], stderr=subprocess.PIPE, text=True)
time.sleep(1.0)
t0 = time.perf_counter()
p.send_signal(signal.SIGINT)
p.communicate()
print(f"exited {time.perf_counter() - t0:.2f}s after SIGINT, rc={p.returncode}")
Use a real process and send_signal for this test: a background job started with & from a non-interactive shell script ignores SIGINT, so kill -INT from such a script tests nothing. Measured: the asyncio version exited 0.03 s after SIGINT — asyncio.run cancelled the main task, which cancelled every in-flight request. The threaded version exited after 4.12 s: the main thread received KeyboardInterrupt immediately, but the interpreter does not exit until worker threads finish, and each was blocked in a socket read that Python cannot interrupt. Both printed a traceback — 38 lines for asyncio, 16 for threads — which is noise for a user who meant to stop the tool.
Verify: press Ctrl-C during the slowest phase of the tool, and measure the time to the prompt and what is printed.
4. Make interruption clean¶
For asyncio, catch KeyboardInterrupt around asyncio.run and exit with the conventional status 130:
def cli():
try:
asyncio.run(main(args.n, args.limit))
except KeyboardInterrupt:
print("interrupted", file=sys.stderr)
sys.exit(130)
Measured: exit in 0.05 s, one line of output, status 130, and the ClientSession closed because cancellation ran its async with exit. For threads, the equivalent — catching the interrupt and calling executor.shutdown(wait=False, cancel_futures=True) — dropped the queued work and printed one line, but the process still took 4.12 s to exit, because the in-flight calls had to finish. The only levers there are short per-call timeouts, which bound the wait, or os._exit, which skips all cleanup. When responsive interruption matters and calls can be slow, that favours asyncio.
Verify: after Ctrl-C, the tool exits within a second, prints one line, returns 130, and leaves no partial output files that look complete.
5. Choose by workload¶
For CPU work, the deciding factor is whether the work releases the GIL. Measured on 32 inputs with 8 workers: SHA-256 over 8 MB buffers took 0.11 s sequentially and 0.02 s with 8 threads, a 4.8× speed-up, because hashlib releases the GIL for large inputs. A pure-Python loop of json.dumps calls took 1.23 s sequentially and 1.27 s with 8 threads — no speed-up — and 0.26 s with an 8-process pool including its start-up, 4.7×.
def process_files(paths, workers=os.cpu_count()):
# Pure-Python per-file work: processes. Hashing or compression: threads are enough.
with ProcessPoolExecutor(workers) as ex:
for path, result in zip(paths, ex.map(analyse, paths, chunksize=8)):
print(path, result)
The resulting rule for a CLI: a handful of calls — stay sequential or use a small thread pool, and keep imports light. Hundreds or thousands of network calls with a need for responsive Ctrl-C — asyncio, with explicit limits at every layer. CPU work in C extensions that release the GIL — threads. Pure-Python CPU work — a process pool. For a tool that mixes them, see combining asyncio with multiprocessing for mixed workloads.
Verify: the chosen model is justified by a measured speed-up on the tool's real input, and interruption has been tested.
Verification¶
The CLI's concurrency is well chosen when:
- Start-up time is measured, and heavy imports are deferred off the
--helppath. - Throughput is measured at the real input size, with every layer's concurrency limit set explicitly.
- Ctrl-C exits within a second with one line of output and status 130.
- CPU work uses threads only where a measured speed-up shows the GIL is released.
Diagnostic Hook: when a tool ignores Ctrl-C for several seconds, check for worker threads blocked in I/O. The time to exit — 4.12 s here — matches the remaining duration of the slowest in-flight call, and adding a timeout to those calls bounds it.
Pitfalls & edge cases¶
- Importing aiohttp for a few requests. Measured: 157 ms of start-up against 34 ms.
- aiohttp's default connector limit. Measured: 1.09 s against 0.62 s with
limit=200. - Blocking calls without timeouts in threads. Ctrl-C waited 4.12 s for them.
- Testing SIGINT with
killfrom a script's background job. Such jobs ignore SIGINT.
Frequently Asked Questions¶
Should a Python CLI tool use asyncio or threads?
For many network calls, both were equally fast here; asyncio exited 0.05 s after Ctrl-C against 4.12 s for threads. For a few calls, threads or sequential code start faster.
Why does my threaded CLI not stop on Ctrl-C?
The interpreter waits for worker threads, and threads blocked in socket reads cannot be interrupted. Exit took 4.12 s, the remaining time of in-flight calls; add timeouts.
How much does importing asyncio slow a CLI's start-up?
A process importing asyncio took 42.4 ms against 9.0 ms empty. aiohttp took 156.5 ms, httpx 89.3 ms and requests 81.7 ms.
What exit code should a CLI use after Ctrl-C?
130, the shell convention for termination by SIGINT. Catch KeyboardInterrupt around asyncio.run, print one line to stderr and call sys.exit(130).
Related¶
- Threading vs Multiprocessing vs asyncio — up to the topic overview.
- asyncio vs threads for database-heavy services — the same question for a long-running service.
- Concurrent Execution & Worker Patterns — the section overview.