Skip to content

Sizing asyncio for Container CPU Limits

Many sizing decisions in an async service derive from the number of CPUs: the default thread-pool size, the number of worker processes, process-pool sizes for CPU-bound work. In a container, the CPU count Python reports is usually the host's, not the container's allowance, so every one of those decisions is made for a much larger machine than the one the service actually gets. Measured with Docker on a 24-core Linux host: in a container started with --cpus=2, Python 3.12 reported os.cpu_count() = 24 and created a default ThreadPoolExecutor with 28 workers; the cgroup file /sys/fs/cgroup/cpu.max said 2.0. With --cpuset-cpus=0-1, os.sched_getaffinity correctly reported 2, but os.cpu_count() still said 24. Python 3.13's os.process_cpu_count() reported 2 under the cpuset — and its executor shrank to 6 workers — but still 24 under the --cpus quota. Setting PYTHON_CPU_COUNT=2 (Python 3.13+) made cpu_count(), process_cpu_count() and the executor all agree with the quota. This guide derives sizes from what the container is actually allowed.

Prerequisites

1. See what Python reports inside the container

Print every source of "how many CPUs" from inside the container:

import concurrent.futures as cf
import os


def cgroup_cpu_quota() -> float | None:
    try:
        quota, period = open("/sys/fs/cgroup/cpu.max").read().split()
    except OSError:
        return None
    return None if quota == "max" else int(quota) / int(period)


print("os.cpu_count()          ", os.cpu_count())
print("os.process_cpu_count()  ", getattr(os, "process_cpu_count", lambda: None)())
print("sched_getaffinity       ", len(os.sched_getaffinity(0)))
print("cgroup cpu.max quota    ", cgroup_cpu_quota())
print("default executor workers", cf.ThreadPoolExecutor()._max_workers)

Measured on Python 3.12 under --cpus=2: 24, (no process_cpu_count), 24, 2.0, 28. Under --cpuset-cpus=0-1: 24, —, 2, max, 28. The two limiting mechanisms differ: a quota (--cpus, Kubernetes CPU limits) caps CPU time across any cores and is visible only in cpu.max; a cpuset pins the process to specific cores and is visible through affinity. Kubernetes CPU limits are quotas, so in most clusters the only accurate source is the cgroup file.

Verify: run the script in your production container image with its production limits, and compare the figures.

What Python reported inside a container on a 24-core host A grid of 6 rows by 5 columns. What Python reported inside a container on a 24-core host container / Python cpu_count process_cpu_count cgroup quota default executor no limit, 3.12 24 n/a max 28 --cpus=2, 3.12 24 n/a 2.0 28 --cpuset-cpus=0-1, 3.12 24 n/a max 28 --cpus=2, 3.13 24 24 2.0 28 --cpuset-cpus=0-1, 3.13 24 2 max 6 --cpus=2, 3.13, PYTHON_CPU_COUNT=2 2 2 2.0 6 Docker on Linux, 24-core host; executor size is min(32, cpus + 4).

2. Derive an effective CPU count from the cgroup

Compute the effective count once, at startup, from the quota and the affinity together, and use it everywhere a CPU count is needed:

import math
import os


def effective_cpus() -> int:
    candidates = [len(os.sched_getaffinity(0))]
    try:
        quota, period = open("/sys/fs/cgroup/cpu.max").read().split()   # cgroup v2
        if quota != "max":
            candidates.append(math.ceil(int(quota) / int(period)))
    except OSError:
        pass
    return max(1, min(candidates))


CPUS = effective_cpus()          # 2 under --cpus=2 and under --cpuset-cpus=0-1

Rounding a fractional quota up — 0.5 becomes 1 — keeps at least one worker while acknowledging that the process cannot use more than half a core on average. On cgroup v1 hosts the equivalent files are cpu.cfs_quota_us and cpu.cfs_period_us; most current distributions and managed Kubernetes services use v2. The PYTHON_CPU_COUNT environment variable, from Python 3.13, is the simplest fix where you control the deployment: set it to the limit, and os.cpu_count(), os.process_cpu_count() and every library that consults them agree.

Verify: effective_cpus() returns the container's limit under both --cpus and --cpuset-cpus.

3. Size the default executor and process pools from it

With an effective count, set the executor explicitly instead of letting Python size it from the host:

from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor


async def main() -> None:
    loop = asyncio.get_running_loop()
    # I/O-bound blocking calls: threads can exceed CPUs, but not by host-sized multiples
    loop.set_default_executor(ThreadPoolExecutor(max_workers=CPUS * 4 + 4,
                                                 thread_name_prefix="default"))
    # CPU-bound work: one process per CPU the container may actually use
    cpu_pool = ProcessPoolExecutor(max_workers=CPUS)

Measured: the host-derived default was 28 threads in a container allowed two CPUs. For blocking I/O that is merely generous; for CPU-bound work submitted to the default executor it means 28 threads competing for two CPUs' worth of quota, which turns into throttling, measured in detecting CPU throttling in async services. A ProcessPoolExecutor() with no argument would start 24 processes, each with its own memory, to share two CPUs. Pool sizing for mixed workloads is covered in optimizing worker pool sizes for mixed I/O and CPU workloads.

Verify: ps -T inside the container shows thread and process counts that match the configured sizes, not host-derived ones.

From container limit to pool sizes A flow of 4 stages. From container limit to pool sizes read cpu.max + affinity quota 2.0, affinity 24 effective CPUs min, rounded up: 2 PYTHON_CPU_COUNT=2 or pass explicitly size pools and workers not from the host's 24 Every derived size should trace back to the container's allowance.

4. Size server workers from the limit, not the node

ASGI servers and process managers default their worker counts from the CPU count too, or are configured from it in deployment manifests:

# Kubernetes: the limit is the source of truth; pass it to the app
resources:
  requests: {cpu: "1", memory: "512Mi"}
  limits:   {cpu: "2", memory: "1Gi"}
env:
  - name: PYTHON_CPU_COUNT
    valueFrom: {resourceFieldRef: {resource: limits.cpu, divisor: "1"}}
  - name: WEB_CONCURRENCY
    valueFrom: {resourceFieldRef: {resource: limits.cpu, divisor: "1"}}

The downward API exposes the container's own CPU limit as an environment variable, so the image needs no cgroup parsing. For an async server, one worker per allowed CPU is a sensible starting point — each worker's event loop uses at most one core — refined by measurement as described in sizing uvicorn workers for async services. Starting 24 workers in a two-CPU container multiplies memory by twelve and leaves them all competing for the same quota.

Verify: the number of server worker processes in each pod equals the CPU limit (or your chosen ratio), checked with ps in a running pod.

5. Report the effective sizes at startup

Make the result visible, so a wrong limit or a missing variable shows up in the first log line rather than in latency graphs:

log.info(
    "cpu sizing: os.cpu_count=%s process_cpu_count=%s affinity=%s quota=%s -> effective=%s; "
    "default executor=%s; process pool=%s",
    os.cpu_count(), getattr(os, "process_cpu_count", lambda: None)(),
    len(os.sched_getaffinity(0)), cgroup_cpu_quota(), CPUS,
    DEFAULT_EXECUTOR._max_workers, CPU_POOL._max_workers,
)

A startup line like this turns the measured discrepancy — 24 reported, 2 allowed — into something an operator notices immediately, and documents which source the service trusted. Export the same values as gauges so dashboards can compare them with observed CPU usage and throttling.

Verify: the startup log of every replica shows an effective CPU count equal to its limit.

Where should this service get its CPU count? A decision on What can you control with 4 outcomes. Where should this service get its CPU count? What can you control? the environment, Python 3.13+ PYTHON_CPU_COUNT=limit fixes every consumer Kubernetes manifests downward API: limits.cpu no parsing in the image only the code read cpu.max + affinity at startup min, rounded up nothing at least log the discrepancy 24 reported, 2 allowed os.cpu_count() describes the host, not the container.

Verification

CPU-derived sizes are correct when:

  • An effective CPU count comes from the cgroup quota and affinity, or from PYTHON_CPU_COUNT.
  • The default executor, process pools and server workers are sized from it.
  • Kubernetes passes the limit to the container through the downward API.
  • Startup logs show reported versus effective counts, and they match the limit.

Diagnostic Hook: compare thread and process counts in a running container with its CPU limit. Twenty-eight default-executor threads or two dozen workers in a two-CPU container is the host's CPU count leaking into sizing, and the throttling metrics in the next guide usually confirm it.

Pitfalls & edge cases

  • Trusting os.cpu_count(). Measured: 24 under a 2-CPU quota.
  • process_cpu_count() and quotas. Measured: it follows cpusets, not --cpus or Kubernetes limits.
  • Unsized ProcessPoolExecutor(). It starts one process per host CPU.
  • Fractional limits. Round up, and remember 0.5 CPU means throttling, not a slower core.

Frequently Asked Questions

Does os.cpu_count() respect Docker CPU limits?

No. In a --cpus=2 container on a 24-core host it returned 24 in testing, and the default ThreadPoolExecutor was sized at 28 workers. Read /sys/fs/cgroup/cpu.max or set PYTHON_CPU_COUNT.

What is os.process_cpu_count() in Python 3.13?

The number of CPUs the process may run on, based on affinity. It returned 2 under --cpuset-cpus=0-1 but 24 under a --cpus=2 quota, so it does not account for Kubernetes-style CPU limits.

How do I make Python see the container's CPU limit?

On Python 3.13+, set PYTHON_CPU_COUNT to the limit (in Kubernetes, from the downward API's limits.cpu); otherwise read the quota from /sys/fs/cgroup/cpu.max and size pools explicitly.

How many uvicorn workers should a container run?

Start from one per CPU the container is allowed — its limit, not the node's core count — and adjust by measurement.