Sizing asyncio for Container CPU Limits¶
Many sizing decisions in an async service derive from the number of CPUs: the default thread-pool size, the number of worker processes, process-pool sizes for CPU-bound work. In a container, the CPU count Python reports is usually the host's, not the container's allowance, so every one of those decisions is made for a much larger machine than the one the service actually gets. Measured with Docker on a 24-core Linux host: in a container started with --cpus=2, Python 3.12 reported os.cpu_count() = 24 and created a default ThreadPoolExecutor with 28 workers; the cgroup file /sys/fs/cgroup/cpu.max said 2.0. With --cpuset-cpus=0-1, os.sched_getaffinity correctly reported 2, but os.cpu_count() still said 24. Python 3.13's os.process_cpu_count() reported 2 under the cpuset — and its executor shrank to 6 workers — but still 24 under the --cpus quota. Setting PYTHON_CPU_COUNT=2 (Python 3.13+) made cpu_count(), process_cpu_count() and the executor all agree with the quota. This guide derives sizes from what the container is actually allowed.
Prerequisites¶
- Docker or Kubernetes; Python 3.11+ (3.13+ for
process_cpu_countandPYTHON_CPU_COUNT). - The default executor, from sizing the default thread pool executor.
- The topic overview, Containers & Serverless.
1. See what Python reports inside the container¶
Print every source of "how many CPUs" from inside the container:
import concurrent.futures as cf
import os
def cgroup_cpu_quota() -> float | None:
try:
quota, period = open("/sys/fs/cgroup/cpu.max").read().split()
except OSError:
return None
return None if quota == "max" else int(quota) / int(period)
print("os.cpu_count() ", os.cpu_count())
print("os.process_cpu_count() ", getattr(os, "process_cpu_count", lambda: None)())
print("sched_getaffinity ", len(os.sched_getaffinity(0)))
print("cgroup cpu.max quota ", cgroup_cpu_quota())
print("default executor workers", cf.ThreadPoolExecutor()._max_workers)
Measured on Python 3.12 under --cpus=2: 24, (no process_cpu_count), 24, 2.0, 28. Under --cpuset-cpus=0-1: 24, —, 2, max, 28. The two limiting mechanisms differ: a quota (--cpus, Kubernetes CPU limits) caps CPU time across any cores and is visible only in cpu.max; a cpuset pins the process to specific cores and is visible through affinity. Kubernetes CPU limits are quotas, so in most clusters the only accurate source is the cgroup file.
Verify: run the script in your production container image with its production limits, and compare the figures.
2. Derive an effective CPU count from the cgroup¶
Compute the effective count once, at startup, from the quota and the affinity together, and use it everywhere a CPU count is needed:
import math
import os
def effective_cpus() -> int:
candidates = [len(os.sched_getaffinity(0))]
try:
quota, period = open("/sys/fs/cgroup/cpu.max").read().split() # cgroup v2
if quota != "max":
candidates.append(math.ceil(int(quota) / int(period)))
except OSError:
pass
return max(1, min(candidates))
CPUS = effective_cpus() # 2 under --cpus=2 and under --cpuset-cpus=0-1
Rounding a fractional quota up — 0.5 becomes 1 — keeps at least one worker while acknowledging that the process cannot use more than half a core on average. On cgroup v1 hosts the equivalent files are cpu.cfs_quota_us and cpu.cfs_period_us; most current distributions and managed Kubernetes services use v2. The PYTHON_CPU_COUNT environment variable, from Python 3.13, is the simplest fix where you control the deployment: set it to the limit, and os.cpu_count(), os.process_cpu_count() and every library that consults them agree.
Verify: effective_cpus() returns the container's limit under both --cpus and --cpuset-cpus.
3. Size the default executor and process pools from it¶
With an effective count, set the executor explicitly instead of letting Python size it from the host:
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor
async def main() -> None:
loop = asyncio.get_running_loop()
# I/O-bound blocking calls: threads can exceed CPUs, but not by host-sized multiples
loop.set_default_executor(ThreadPoolExecutor(max_workers=CPUS * 4 + 4,
thread_name_prefix="default"))
# CPU-bound work: one process per CPU the container may actually use
cpu_pool = ProcessPoolExecutor(max_workers=CPUS)
Measured: the host-derived default was 28 threads in a container allowed two CPUs. For blocking I/O that is merely generous; for CPU-bound work submitted to the default executor it means 28 threads competing for two CPUs' worth of quota, which turns into throttling, measured in detecting CPU throttling in async services. A ProcessPoolExecutor() with no argument would start 24 processes, each with its own memory, to share two CPUs. Pool sizing for mixed workloads is covered in optimizing worker pool sizes for mixed I/O and CPU workloads.
Verify: ps -T inside the container shows thread and process counts that match the configured sizes, not host-derived ones.
4. Size server workers from the limit, not the node¶
ASGI servers and process managers default their worker counts from the CPU count too, or are configured from it in deployment manifests:
# Kubernetes: the limit is the source of truth; pass it to the app
resources:
requests: {cpu: "1", memory: "512Mi"}
limits: {cpu: "2", memory: "1Gi"}
env:
- name: PYTHON_CPU_COUNT
valueFrom: {resourceFieldRef: {resource: limits.cpu, divisor: "1"}}
- name: WEB_CONCURRENCY
valueFrom: {resourceFieldRef: {resource: limits.cpu, divisor: "1"}}
The downward API exposes the container's own CPU limit as an environment variable, so the image needs no cgroup parsing. For an async server, one worker per allowed CPU is a sensible starting point — each worker's event loop uses at most one core — refined by measurement as described in sizing uvicorn workers for async services. Starting 24 workers in a two-CPU container multiplies memory by twelve and leaves them all competing for the same quota.
Verify: the number of server worker processes in each pod equals the CPU limit (or your chosen ratio), checked with ps in a running pod.
5. Report the effective sizes at startup¶
Make the result visible, so a wrong limit or a missing variable shows up in the first log line rather than in latency graphs:
log.info(
"cpu sizing: os.cpu_count=%s process_cpu_count=%s affinity=%s quota=%s -> effective=%s; "
"default executor=%s; process pool=%s",
os.cpu_count(), getattr(os, "process_cpu_count", lambda: None)(),
len(os.sched_getaffinity(0)), cgroup_cpu_quota(), CPUS,
DEFAULT_EXECUTOR._max_workers, CPU_POOL._max_workers,
)
A startup line like this turns the measured discrepancy — 24 reported, 2 allowed — into something an operator notices immediately, and documents which source the service trusted. Export the same values as gauges so dashboards can compare them with observed CPU usage and throttling.
Verify: the startup log of every replica shows an effective CPU count equal to its limit.
Verification¶
CPU-derived sizes are correct when:
- An effective CPU count comes from the cgroup quota and affinity, or from
PYTHON_CPU_COUNT. - The default executor, process pools and server workers are sized from it.
- Kubernetes passes the limit to the container through the downward API.
- Startup logs show reported versus effective counts, and they match the limit.
Diagnostic Hook: compare thread and process counts in a running container with its CPU limit. Twenty-eight default-executor threads or two dozen workers in a two-CPU container is the host's CPU count leaking into sizing, and the throttling metrics in the next guide usually confirm it.
Pitfalls & edge cases¶
- Trusting
os.cpu_count(). Measured: 24 under a 2-CPU quota. process_cpu_count()and quotas. Measured: it follows cpusets, not--cpusor Kubernetes limits.- Unsized
ProcessPoolExecutor(). It starts one process per host CPU. - Fractional limits. Round up, and remember 0.5 CPU means throttling, not a slower core.
Frequently Asked Questions¶
Does os.cpu_count() respect Docker CPU limits?
No. In a --cpus=2 container on a 24-core host it returned 24 in testing, and the default ThreadPoolExecutor was sized at 28 workers. Read /sys/fs/cgroup/cpu.max or set PYTHON_CPU_COUNT.
What is os.process_cpu_count() in Python 3.13?
The number of CPUs the process may run on, based on affinity. It returned 2 under --cpuset-cpus=0-1 but 24 under a --cpus=2 quota, so it does not account for Kubernetes-style CPU limits.
How do I make Python see the container's CPU limit?
On Python 3.13+, set PYTHON_CPU_COUNT to the limit (in Kubernetes, from the downward API's limits.cpu); otherwise read the quota from /sys/fs/cgroup/cpu.max and size pools explicitly.
How many uvicorn workers should a container run?
Start from one per CPU the container is allowed — its limit, not the node's core count — and adjust by measurement.
Related¶
- Containers & Serverless — up to the topic overview.
- Detecting CPU throttling in async services — what happens when sizes exceed the quota.
- Resilience, Cancellation & Error Handling — the section overview.