Comparing Memory per asyncio Task, Thread and Process¶
"asyncio scales to more connections than threads" is usually said without numbers, and the numbers decide architecture: how many concurrent connections one container can hold, whether a thread per connection is viable, how many worker processes fit. Measured on Python 3.14 on Linux, an idle asyncio task holding a 1 KiB buffer cost 2.2 KiB of resident memory; an idle thread holding the same buffer cost 19.2 KiB of resident memory, while its virtual size grew by about 72 MiB (an 8 MiB stack reservation plus a glibc malloc arena); and a bare child process — no application imports — cost 11.0 to 15.5 MiB of RSS depending on the start method, of which 1.8 to 7.7 MiB was its proportional share once memory shared with the parent was accounted for. At 10,000 concurrent connections that is roughly 22 MiB for tasks, 190 MiB for threads, and not a realistic option for processes. This guide shows how to take these measurements for your own workload, because your per-unit cost is dominated by what each one holds.
Prerequisites¶
- Python 3.11+ on Linux (the measurements read
/proc), stdlib only. - The model comparison, from Threading vs Multiprocessing vs Asyncio.
- Memory tooling, from finding memory leaks in asyncio with tracemalloc.
1. Measure tasks by RSS delta¶
Create many idle units, each holding what one connection would hold, and compare resident memory before and after:
import asyncio
def rss_kib() -> int:
with open("/proc/self/status") as f:
for line in f:
if line.startswith("VmRSS"):
return int(line.split()[1])
return 0
async def per_task(n: int = 10_000) -> float:
before = rss_kib()
release = asyncio.Event()
async def idle():
buf = bytearray(1024) # what a connection handler keeps alive
await release.wait()
tasks = [asyncio.create_task(idle()) for _ in range(n)]
await asyncio.sleep(0.2)
after = rss_kib()
release.set()
await asyncio.gather(*tasks)
return (after - before) / n
print(asyncio.run(per_task())) # ~2.2 KiB on Python 3.14
The 2.2 KiB is the task object, its coroutine frame, the event waiter, and the 1 KiB buffer. Real handlers hold more — a parsed request, a database row, response chunks — and that dominates quickly: a handler with 50 KiB of locals costs about 51 KiB per connection regardless of the concurrency model. Use the same harness with your real handler's state.
Verify: run with n = 1,000 and 10,000; the per-unit figure should be stable, confirming you are measuring per-task cost and not fixed overhead.
2. Measure threads, and know what RSS hides¶
The same harness with threads:
import threading
import time
def per_thread(n: int = 1_000) -> float:
before = rss_kib()
release = threading.Event()
def idle():
buf = bytearray(1024)
release.wait()
threads = [threading.Thread(target=idle) for _ in range(n)]
for t in threads:
t.start()
time.sleep(0.5)
after = rss_kib()
release.set()
for t in threads:
t.join()
return (after - before) / n # ~19 KiB resident each
Resident cost is modest — 19.2 KiB — because a thread's stack is reserved, not committed: Linux reserves 8 MiB of address space per thread by default and only backs the pages actually touched. The virtual size tells a different story: creating 100 idle threads grew VmSize by about 72 MiB per thread, the 8 MiB stack plus a 64 MiB malloc arena that glibc reserves for new threads (up to a default cap of eight arenas per CPU). Two consequences. Deep recursion or large stack frames make the resident figure grow. And address space can run out before RAM does where ulimit -v or a sandbox counts it. threading.stack_size() shrinks stack reservations, and MALLOC_ARENA_MAX caps the arenas, for thread-heavy designs.
Verify: compare VmSize (virtual) as well as VmRSS before and after creating threads; expect tens of megabytes of virtual growth per thread for the first threads.
3. Measure processes with PSS, not RSS¶
For processes, RSS double-counts pages shared with the parent. The proportional set size (PSS) divides shared pages among their users and gives the fair cost:
import multiprocessing as mp
def pss_kib(pid: int) -> int:
with open(f"/proc/{pid}/smaps_rollup") as f:
for line in f:
if line.startswith("Pss:"):
return int(line.split()[1])
return 0
def child(ev):
ev.wait()
if __name__ == "__main__":
for method in ("fork", "forkserver", "spawn"):
ctx = mp.get_context(method)
ev = ctx.Event()
procs = [ctx.Process(target=child, args=(ev,)) for _ in range(8)]
for p in procs:
p.start()
... # sleep, then sum pss_kib(p.pid) and RSS for each
Measured for bare children: fork 11.0 MiB RSS / 1.8 MiB PSS, forkserver 13.3 / 3.1, spawn 15.5 / 7.7. Fork shares the most with the parent at first; spawn shares the least because it starts a fresh interpreter. These are floors: a worker that imports a web framework, an ORM and a few SDKs typically costs tens of megabytes before handling a request, and copy-on-write sharing after fork erodes as reference counts touch shared pages. Measure your real worker's PSS after warm-up, as in sizing uvicorn workers for async services.
Verify: your service's worker PSS after handling traffic for a few minutes is the number to use for capacity planning.
4. Turn per-unit costs into capacity¶
With per-unit costs measured, capacity is arithmetic. For 10,000 concurrent idle connections each holding about 1 KiB:
| Model | Per unit | 10,000 units | Notes |
|---|---|---|---|
| asyncio tasks | 2.2 KiB | ~22 MiB | one thread, one loop |
| threads | 19.2 KiB RSS | ~190 MiB RSS, tens of GiB virtual | stacks + malloc arenas, scheduler load |
| processes | 1.8–15.5 MiB | 18–155 GiB | not viable per connection |
The architectural conclusion is the familiar one, now with numbers: per-connection processes are out of the question at this scale, per-connection threads are feasible into the low thousands, and tasks scale to tens of thousands on memory alone. In practice, CPU for parsing and handling — not memory for idle connections — is what limits an asyncio server, which is why production services combine one loop per process with a handful of processes, as discussed in combining asyncio with multiprocessing for mixed workloads.
Verify: plug your measured per-connection cost into the table and compare with your container limit at peak connection count.
5. Watch the per-unit cost in production¶
The per-unit cost drifts as code changes: a handler that starts keeping the full request body, a middleware that caches per-connection state. Track it directly:
async def report_memory_per_connection(get_connections, every: float = 60.0) -> None:
while True:
conns = max(1, get_connections())
metrics.gauge("rss_kib", rss_kib())
metrics.gauge("open_connections", conns)
metrics.gauge("rss_kib_per_connection", rss_kib() / conns)
await asyncio.sleep(every)
RSS per connection at steady state is the empirical version of step 1. If it rises across releases with no change in traffic shape, a handler is holding more per request; if it rises over the life of one process, that is a leak, and the tools in tracking task growth in long-running services apply.
Verify: the gauge is stable across a day of traffic for a given release.
Verification¶
Your capacity numbers are trustworthy when:
- Per-unit cost is measured with your real handler state, not just an empty coroutine.
- Threads are evaluated on virtual as well as resident memory.
- Processes are measured by PSS after warm-up, not RSS at start.
- RSS per connection is tracked in production and stable across a release.
Diagnostic Hook: alert when RSS per open connection deviates by more than 25% from the previous release's baseline at similar traffic. It is a cheap, high-signal check that catches handlers starting to retain request bodies, caches attached to connections, and leaks — long before the container hits its memory limit.
Pitfalls & edge cases¶
- Measuring empty coroutines. The primitive is cheap; what each handler holds is not.
- Reading RSS for forked processes. Shared pages are counted once per process; use PSS.
- Ignoring virtual memory for threads. Address-space limits can fail thread creation long before RAM runs out.
- Assuming fork sharing persists. Reference-count writes copy shared pages as workers run.
Frequently Asked Questions¶
How much memory does an asyncio task use?
An idle task holding a 1 KiB buffer measured 2.2 KiB of resident memory on Python 3.14. Real cost is dominated by what the handler keeps alive in its frame.
How much memory does a Python thread use?
An idle thread measured about 19 KiB of resident memory. Its virtual size grew by about 72 MiB: an 8 MiB stack reservation plus a glibc malloc arena. Deep stacks raise the resident cost.
How much memory does a Python worker process use?
A bare child measured 11 to 15.5 MiB RSS depending on start method, but only 1.8 to 7.7 MiB proportional set size once shared pages are divided. Application imports usually add tens of megabytes.
Can I run a thread per connection instead of asyncio?
For a few hundred connections, yes. At 10,000 connections threads cost around 190 MiB resident and tens of gigabytes of reserved address space, against about 22 MiB for asyncio tasks.
Related¶
- Threading vs Multiprocessing vs Asyncio — up to the topic overview.
- When asyncio is slower than threads — the other side of the trade-off.
- Concurrent Execution & Worker Patterns — the section overview.