Skip to content

Comparing Memory per asyncio Task, Thread and Process

"asyncio scales to more connections than threads" is usually said without numbers, and the numbers decide architecture: how many concurrent connections one container can hold, whether a thread per connection is viable, how many worker processes fit. Measured on Python 3.14 on Linux, an idle asyncio task holding a 1 KiB buffer cost 2.2 KiB of resident memory; an idle thread holding the same buffer cost 19.2 KiB of resident memory, while its virtual size grew by about 72 MiB (an 8 MiB stack reservation plus a glibc malloc arena); and a bare child process — no application imports — cost 11.0 to 15.5 MiB of RSS depending on the start method, of which 1.8 to 7.7 MiB was its proportional share once memory shared with the parent was accounted for. At 10,000 concurrent connections that is roughly 22 MiB for tasks, 190 MiB for threads, and not a realistic option for processes. This guide shows how to take these measurements for your own workload, because your per-unit cost is dominated by what each one holds.

Prerequisites

1. Measure tasks by RSS delta

Create many idle units, each holding what one connection would hold, and compare resident memory before and after:

import asyncio


def rss_kib() -> int:
    with open("/proc/self/status") as f:
        for line in f:
            if line.startswith("VmRSS"):
                return int(line.split()[1])
    return 0


async def per_task(n: int = 10_000) -> float:
    before = rss_kib()
    release = asyncio.Event()

    async def idle():
        buf = bytearray(1024)          # what a connection handler keeps alive
        await release.wait()

    tasks = [asyncio.create_task(idle()) for _ in range(n)]
    await asyncio.sleep(0.2)
    after = rss_kib()
    release.set()
    await asyncio.gather(*tasks)
    return (after - before) / n


print(asyncio.run(per_task()))       # ~2.2 KiB on Python 3.14

The 2.2 KiB is the task object, its coroutine frame, the event waiter, and the 1 KiB buffer. Real handlers hold more — a parsed request, a database row, response chunks — and that dominates quickly: a handler with 50 KiB of locals costs about 51 KiB per connection regardless of the concurrency model. Use the same harness with your real handler's state.

Verify: run with n = 1,000 and 10,000; the per-unit figure should be stable, confirming you are measuring per-task cost and not fixed overhead.

Resident memory per idle unit of concurrency 4 horizontal bars comparing process, spawn with the others. Resident memory per idle unit of concurrency process, spawn 15.5 MiB RSS process, fork 11.0 MiB RSS thread 19.2 KiB RSS asyncio task 2.2 KiB RSS Each unit holds a 1 KiB buffer and waits; processes import nothing beyond multiprocessing. Processes are three orders of magnitude heavier than threads, threads about nine times heavier than tasks.

2. Measure threads, and know what RSS hides

The same harness with threads:

import threading
import time


def per_thread(n: int = 1_000) -> float:
    before = rss_kib()
    release = threading.Event()

    def idle():
        buf = bytearray(1024)
        release.wait()

    threads = [threading.Thread(target=idle) for _ in range(n)]
    for t in threads:
        t.start()
    time.sleep(0.5)
    after = rss_kib()
    release.set()
    for t in threads:
        t.join()
    return (after - before) / n            # ~19 KiB resident each

Resident cost is modest — 19.2 KiB — because a thread's stack is reserved, not committed: Linux reserves 8 MiB of address space per thread by default and only backs the pages actually touched. The virtual size tells a different story: creating 100 idle threads grew VmSize by about 72 MiB per thread, the 8 MiB stack plus a 64 MiB malloc arena that glibc reserves for new threads (up to a default cap of eight arenas per CPU). Two consequences. Deep recursion or large stack frames make the resident figure grow. And address space can run out before RAM does where ulimit -v or a sandbox counts it. threading.stack_size() shrinks stack reservations, and MALLOC_ARENA_MAX caps the arenas, for thread-heavy designs.

Verify: compare VmSize (virtual) as well as VmRSS before and after creating threads; expect tens of megabytes of virtual growth per thread for the first threads.

3. Measure processes with PSS, not RSS

For processes, RSS double-counts pages shared with the parent. The proportional set size (PSS) divides shared pages among their users and gives the fair cost:

import multiprocessing as mp


def pss_kib(pid: int) -> int:
    with open(f"/proc/{pid}/smaps_rollup") as f:
        for line in f:
            if line.startswith("Pss:"):
                return int(line.split()[1])
    return 0


def child(ev):
    ev.wait()


if __name__ == "__main__":
    for method in ("fork", "forkserver", "spawn"):
        ctx = mp.get_context(method)
        ev = ctx.Event()
        procs = [ctx.Process(target=child, args=(ev,)) for _ in range(8)]
        for p in procs:
            p.start()
        ...  # sleep, then sum pss_kib(p.pid) and RSS for each

Measured for bare children: fork 11.0 MiB RSS / 1.8 MiB PSS, forkserver 13.3 / 3.1, spawn 15.5 / 7.7. Fork shares the most with the parent at first; spawn shares the least because it starts a fresh interpreter. These are floors: a worker that imports a web framework, an ORM and a few SDKs typically costs tens of megabytes before handling a request, and copy-on-write sharing after fork erodes as reference counts touch shared pages. Measure your real worker's PSS after warm-up, as in sizing uvicorn workers for async services.

Verify: your service's worker PSS after handling traffic for a few minutes is the number to use for capacity planning.

Bare child process cost by start method A grid of 3 rows by 4 columns. Bare child process cost by start method start method RSS per process PSS per process default on fork 11.0 MiB 1.8 MiB Linux up to 3.13 forkserver 13.3 MiB 3.1 MiB Linux 3.14+ spawn 15.5 MiB 7.7 MiB macOS, Windows PSS is the honest per-process number; RSS counts shared pages once per process.

4. Turn per-unit costs into capacity

With per-unit costs measured, capacity is arithmetic. For 10,000 concurrent idle connections each holding about 1 KiB:

Model Per unit 10,000 units Notes
asyncio tasks 2.2 KiB ~22 MiB one thread, one loop
threads 19.2 KiB RSS ~190 MiB RSS, tens of GiB virtual stacks + malloc arenas, scheduler load
processes 1.8–15.5 MiB 18–155 GiB not viable per connection

The architectural conclusion is the familiar one, now with numbers: per-connection processes are out of the question at this scale, per-connection threads are feasible into the low thousands, and tasks scale to tens of thousands on memory alone. In practice, CPU for parsing and handling — not memory for idle connections — is what limits an asyncio server, which is why production services combine one loop per process with a handful of processes, as discussed in combining asyncio with multiprocessing for mixed workloads.

Verify: plug your measured per-connection cost into the table and compare with your container limit at peak connection count.

5. Watch the per-unit cost in production

The per-unit cost drifts as code changes: a handler that starts keeping the full request body, a middleware that caches per-connection state. Track it directly:

async def report_memory_per_connection(get_connections, every: float = 60.0) -> None:
    while True:
        conns = max(1, get_connections())
        metrics.gauge("rss_kib", rss_kib())
        metrics.gauge("open_connections", conns)
        metrics.gauge("rss_kib_per_connection", rss_kib() / conns)
        await asyncio.sleep(every)

RSS per connection at steady state is the empirical version of step 1. If it rises across releases with no change in traffic shape, a handler is holding more per request; if it rises over the life of one process, that is a leak, and the tools in tracking task growth in long-running services apply.

Verify: the gauge is stable across a day of traffic for a given release.

Which unit of concurrency per connection? A decision on How many concurrent connections per worker with 3 outcomes. Which unit of concurrency per connection? How many concurrent connections per worker? up to a few hundred threads are fine 19 KiB each thousands and more asyncio tasks 2.2 KiB each any number processes as a pool never one per connection Memory per unit decides the ceiling; measure your handler's state, not just the primitive.

Verification

Your capacity numbers are trustworthy when:

  • Per-unit cost is measured with your real handler state, not just an empty coroutine.
  • Threads are evaluated on virtual as well as resident memory.
  • Processes are measured by PSS after warm-up, not RSS at start.
  • RSS per connection is tracked in production and stable across a release.

Diagnostic Hook: alert when RSS per open connection deviates by more than 25% from the previous release's baseline at similar traffic. It is a cheap, high-signal check that catches handlers starting to retain request bodies, caches attached to connections, and leaks — long before the container hits its memory limit.

Pitfalls & edge cases

  • Measuring empty coroutines. The primitive is cheap; what each handler holds is not.
  • Reading RSS for forked processes. Shared pages are counted once per process; use PSS.
  • Ignoring virtual memory for threads. Address-space limits can fail thread creation long before RAM runs out.
  • Assuming fork sharing persists. Reference-count writes copy shared pages as workers run.

Frequently Asked Questions

How much memory does an asyncio task use?

An idle task holding a 1 KiB buffer measured 2.2 KiB of resident memory on Python 3.14. Real cost is dominated by what the handler keeps alive in its frame.

How much memory does a Python thread use?

An idle thread measured about 19 KiB of resident memory. Its virtual size grew by about 72 MiB: an 8 MiB stack reservation plus a glibc malloc arena. Deep stacks raise the resident cost.

How much memory does a Python worker process use?

A bare child measured 11 to 15.5 MiB RSS depending on start method, but only 1.8 to 7.7 MiB proportional set size once shared pages are divided. Application imports usually add tens of megabytes.

Can I run a thread per connection instead of asyncio?

For a few hundred connections, yes. At 10,000 connections threads cost around 190 MiB resident and tens of gigabytes of reserved address space, against about 22 MiB for asyncio tasks.