Skip to content

Combining asyncio with Multiprocessing for Mixed Workloads

Most real services are neither purely I/O-bound nor purely CPU-bound: a request waits on a database or an API, then spends a few milliseconds transforming, rendering or validating. asyncio handles the waiting; the GIL caps the computing at one core per process. Two topologies scale past that cap — one event loop feeding a process pool, or several independent processes each running its own loop — and they reach almost the same throughput. Measured with requests that each awaited 20 ms of simulated I/O and then did 7.6 ms of pure-Python CPU work, one process managed 129 requests per second; one loop with a process pool of 4 managed 479, and with 8, 871; four independent loop processes managed 495, and eight, 873. Since throughput does not separate them, this guide compares the two on what does: latency of light requests, memory, connection reuse and operational simplicity.

Prerequisites

1. Measure the ceiling of one process first

With CPU work inline, one process completes at most about 1 / cpu_per_request requests per second, no matter how much I/O concurrency the loop provides:

import asyncio
import time


def cpu(_=None) -> int:                      # ~7.6 ms of pure-Python work
    s = 0
    for i in range(400_000):
        s += i
    return s


async def serve(n: int, pool=None, concurrency: int = 100) -> float:
    sem = asyncio.Semaphore(concurrency)
    loop = asyncio.get_running_loop()

    async def request():
        async with sem:
            await asyncio.sleep(0.02)                      # the I/O part
            if pool is None:
                cpu()                                      # inline: blocks the loop
            else:
                await loop.run_in_executor(pool, cpu)      # offloaded
    t = time.perf_counter()
    await asyncio.gather(*(request() for _ in range(n)))
    return n / (time.perf_counter() - t)

Inline: 129 requests per second, close to the 1 / 7.6 ms ≈ 131 ceiling. The I/O concurrency was fully used — 100 requests waiting at once — and made no difference, because the loop thread was busy computing for almost all of each second. That number is the starting point for any scaling decision.

Verify: measure CPU per request in your service (time.process_time() deltas) and compare 1 / that number with observed throughput per process.

Requests per second with 20 ms I/O and 7.6 ms CPU each 5 horizontal bars comparing 8 loop processes, CPU inline with the others. Requests per second with 20 ms I/O and 7.6 ms CPU each 8 loop processes, CPU inline 873 req/s 1 loop + pool of 8 871 req/s 4 loop processes, CPU inline 495 req/s 1 loop + pool of 4 479 req/s 1 process, CPU inline 129 req/s Python 3.14 on 24 cores; concurrency 100 per loop. Both topologies scale almost linearly with process count; throughput does not choose between them.

2. Topology A: one loop, a process pool for CPU

from concurrent.futures import ProcessPoolExecutor

cpu_pool = ProcessPoolExecutor(max_workers=8)


async def handle(request):
    data = await fetch_from_db(request.id)                       # loop: I/O
    loop = asyncio.get_running_loop()
    result = await loop.run_in_executor(cpu_pool, transform, data)   # pool: CPU
    return JSONResponse(result)

Strengths: one event loop holds every connection, so there is one HTTP client pool, one database pool and one in-memory cache — connection reuse is maximal and state is shared without coordination. Requests that need no CPU never queue behind CPU work, because the loop is free while the pool computes. Costs: every CPU step pickles its input and output across a process boundary, so it suits steps with modest data and real computation; and one loop thread still does all the I/O bookkeeping, which becomes the bottleneck at very high request rates.

Verify: loop lag stays low while the pool is saturated, and light endpoints keep their latency.

3. Topology B: N processes, each with its own loop

This is what uvicorn --workers 8 or gunicorn -k uvicorn.workers.UvicornWorker -w 8 gives you: independent processes behind one listening socket, each a complete asyncio application with CPU work inline.

gunicorn app:app -k uvicorn.workers.UvicornWorker -w 8 --bind 0.0.0.0:8000

Strengths: no pickling — data never crosses a process boundary inside a request — and no special code; the application is written as if it were single-process. The loop-thread bottleneck is divided by N. Costs: every process has its own connection pools (eight database pools of 20 is 160 connections), its own caches (eight cold caches after a deploy), and its own memory footprint. And CPU work still blocks the loop of the process it runs in, so light requests that land on a process mid-computation wait — the tail-latency effect from how the GIL affects async services, divided across processes rather than eliminated.

Verify: total database connections equal workers × pool size and stay under the database's limit.

Loop plus pool versus N loop processes A grid of 5 rows by 3 columns. Loop plus pool versus N loop processes property 1 loop + process pool N loop processes throughput, 8 processes 871 req/s 873 req/s light requests during CPU work unaffected wait if same process data crossing processes every CPU step none connection pools and caches one shared set N separate sets code changes offload calls none Same throughput; the choice is about tail latency, pickling and how many pools you run.

4. Combine them when both matter

The topologies compose. A common production shape is a few worker processes for availability and loop capacity, each with a small process pool for its heaviest step:

# each of 4 uvicorn workers creates its own small pool at startup
@asynccontextmanager
async def lifespan(app):
    app.state.cpu_pool = ProcessPoolExecutor(max_workers=2)
    yield
    app.state.cpu_pool.shutdown(wait=True, cancel_futures=True)


async def handle(request):
    data = await fetch(request)
    if len(data) > 10_000:                                   # only big inputs pay the hop
        return await asyncio.get_running_loop().run_in_executor(
            request.app.state.cpu_pool, transform, data)
    return transform(data)                                    # small inputs inline

Size the total — workers × (1 + pool size) — against the CPU count, leaving headroom. Offloading only above a size threshold keeps small requests fast (no pickling) while protecting the loop from the large ones. Pool lifecycle in an ASGI app is covered in managing startup and shutdown with ASGI lifespan.

Verify: under mixed traffic, p99 for small requests stays near their unloaded value and CPU utilisation is spread across all processes.

5. Choose from three measurements

Decide with numbers rather than preference:

def recommend(cpu_ms_per_req: float, payload_kib: float, pickle_ms: float,
              light_p99_budget_ms: float) -> str:
    if cpu_ms_per_req < 1:
        return "N loop processes, CPU inline: offloading costs more than it saves"
    if pickle_ms > cpu_ms_per_req / 2:
        return "N loop processes: data transfer would eat the gain"
    if cpu_ms_per_req > light_p99_budget_ms / 2:
        return "process pool for the heavy step: inline CPU breaks the light p99 budget"
    return "either; prefer N processes for simplicity"

The inputs are CPU per request, the cost of pickling the step's input and output (measure with pickle.dumps timing on real data), and the tail-latency budget for requests that do no CPU work. Most CRUD-style services land in "N processes"; services with a few heavy endpoints among many light ones land in "process pool for the heavy step".

Verify: the chosen topology is documented with the three measurements that justified it.

Which topology for this mixed workload? A decision on What does CPU per request do to light requests with 3 outcomes. Which topology for this mixed workload? What does CPU per request do to light requests? little: under ~1 ms N loop processes no pickling breaks their p99 pool for heavy steps loop stays free both, at scale workers + small pools offload above a size Throughput scales either way; pick by tail latency and transfer cost.

Verification

The topology fits when:

  • Throughput per process matches 1 / CPU-per-request when CPU is inline, or scales with pool size when offloaded.
  • Light requests keep their p99 during CPU-heavy traffic.
  • Pickling cost is a small fraction of offloaded CPU time.
  • Total connections across processes stay within downstream limits.

Diagnostic Hook: export per-process CPU utilisation, loop lag and, for pools, queue depth and pickled bytes per call. A pool with a growing queue needs more workers; high loop lag with low pool usage means CPU work is still running inline somewhere; pickled bytes per call rising faster than CPU time means the offloaded step should move to shared memory or back inline.

Pitfalls & edge cases

  • Offloading tiny steps. The hop costs more than the work.
  • N workers × large pools. Connection limits are exhausted at deploy time.
  • Process pools created before forking workers. Create pools inside each worker's lifespan.
  • Counting cores twice. Workers plus pool processes must fit the CPU count with headroom.

Frequently Asked Questions

How do I scale an asyncio service that also does CPU work?

Either run several worker processes, each with its own event loop, or keep one loop and send CPU-heavy steps to a process pool. Both scaled from 129 to about 870 requests per second with 8 processes in testing.

Is a process pool better than more uvicorn workers?

Not for throughput, which was nearly identical. A pool protects light requests' latency from CPU work and shares connection pools; more workers avoid pickling and need no code changes.

How many processes should I run?

Enough that per-process CPU stays below about 70% at peak, counting both workers and pool processes against the core count, and few enough that per-process connection pools fit the database's limit.

When is offloading CPU work to a process not worth it?

When the work takes under about a millisecond, or when pickling its input and output costs a large fraction of the work itself.