Async in Containers & Serverless¶
Most asyncio services run inside a container or a serverless runtime, and both environments change rules that Python programs normally take for granted. The process may be PID 1, where unhandled signals are ignored. The CPU count and memory size the interpreter reports describe the host, not the allowance. CPU limits are enforced by stopping the process for parts of every 100 ms, and memory limits by killing it. In serverless, the event loop has to be bridged into a synchronous handler and survive freezes between invocations. This section measures each of these on Docker and the AWS Lambda Python base image. An asyncio program as PID 1 with no SIGTERM handler took 10.26 s to stop and was killed with exit 137; with a handler, 0.44 s and exit 0. In a --cpus=2 container, Python reported 24 CPUs and sized its default executor at 28 threads. Under a 0.5-CPU quota, a busy background thread pushed event-loop lag p99 from 10.6 to 59.6 ms, with 2.6 s of throttling in five seconds. In a 256 MiB container Python reported 60.9 GiB of memory, and an unbounded queue was OOM-killed in 0.9 s with no cleanup. In Lambda, asyncio.run per invocation took 0.87 ms warm against 0.19 ms with a module-level loop — and a global session reused across asyncio.run calls failed 20 of 20 invocations.
The parent section, Resilience, Cancellation & Error Handling, covers shutdown, timeouts and load shedding in general; this topic covers how container and serverless runtimes change them, and how to tell from inside a running service which of their rules is causing a problem.
Architectural principles¶
- Handle
SIGTERMexplicitly, and make sure Python is the process that receives it. - Derive every size from the container's limits — CPU quota, affinity, memory — never from host-reported figures.
- Treat throttling as a latency source and measure it directly from the cgroup.
- Stay under the memory limit by construction, because the OOM killer runs no cleanup.
- In serverless, the loop's lifetime decides what can be reused, and nothing may run between invocations.
Execution model: the runtime around the event loop¶
Inside a container, the event loop runs exactly as it does anywhere else; what differs is the environment's response to it. Signals: Docker and Kubernetes stop a container by signalling its PID 1. If Python is PID 1 and has no SIGTERM handler, the kernel ignores the signal — the PID-1 rule — and the platform kills the process after the grace period. If a shell is PID 1, the signal never reaches Python at all. CPU: a limit of 0.5 CPU means 50 ms of CPU time per 100 ms period for all of the process's threads together. A burst that uses the budget early stops every thread, including the one running the loop, until the period ends — invisible in average utilization, obvious in loop lag. Memory: the cgroup charges the container's usage, including page cache, against a hard limit; crossing it brings the OOM killer's SIGKILL.
Lambda adds an invocation model on top. The runtime calls a synchronous handler once per invocation; the execution environment, with its module-level state, is reused for later invocations and frozen between them. An event loop created by asyncio.run in the handler dies at the end of each invocation, taking with it anything bound to it; a loop created at module level lives as long as the environment. Nothing runs while frozen — not background tasks, not keep-alive timers — while servers on the other end of pooled connections keep counting their idle timeouts.
Pattern catalogue¶
A SIGTERM handler, with Python as PID 1 (or behind an init)¶
loop.add_signal_handler(signal.SIGTERM, asyncio.current_task().cancel)
Exec-form CMD, exec at the end of entrypoint scripts, and tini when the application spawns children. The measurements separate the pieces: the handler alone fixed the stop (0.44 s, exit 0); --init alone made it fast but abrupt (0.22 s, exit 143, no cleanup); a shell-form command defeated even a correct handler (10.25 s, exit 137). See handling signals as PID 1 in containers.
Sizes from the cgroup¶
Effective CPUs from cpu.max and affinity (or PYTHON_CPU_COUNT on 3.13+), and pools and workers sized from them. Python 3.13's os.process_cpu_count() followed a cpuset (2 CPUs, 6 executor threads) but not a quota, which is how Kubernetes limits work, so the cgroup file or the environment variable is the reliable source. See sizing asyncio for container CPU limits.
Throttling as a metric¶
nr_throttled and throttled_usec from cpu.stat, plotted against loop lag; CPU work moved out of latency-critical processes. The same busy thread produced 10.6 ms of lag p99 without a limit (GIL contention, zero throttling) and 59.6 ms under 0.5 CPU (51 throttled periods in five seconds), so the counters, not the lag alone, identify the cause. See detecting CPU throttling in async services.
Memory bounds and a guard¶
Queues and concurrency sized from memory.max, and a guard that sheds load at 80% of it. The guard turned a 0.9-second OOM kill with no cleanup into a controlled stop at 212 MiB with the finally block run and an exit status the program chose. See respecting container memory limits in async services.
A module-level loop in Lambda¶
LOOP = asyncio.new_event_loop()
def handler(event, context):
return LOOP.run_until_complete(work(event))
The module-level loop is what lets a session and its pooled connections survive from one invocation to the next: 0.19 ms per warm invocation against 0.87 ms for asyncio.run with a fresh session. After a 4-second freeze past the upstream's 2-second keep-alive timeout, aiohttp replaced the dead connection transparently (2.35 ms) — safe for reads; writes should retry once with an idempotency key. See running asyncio in AWS Lambda handlers and reusing connections across Lambda invocations.
Reproducing the runtime locally¶
Every behaviour in this section can be reproduced on a laptop with Docker, which makes them testable before they appear in production. The flags map directly onto the production mechanisms:
# PID 1 and signals: time the stop and check the exit code
docker run -d --name t app:latest && sleep 3 && time docker stop -t 10 t
docker inspect -f '{{.State.ExitCode}} OOMKilled={{.State.OOMKilled}}' t
# CPU quota and cpuset, as Kubernetes limits and pinned nodes would apply them
docker run --rm --cpus=0.5 app:latest python -m app.selfcheck
docker run --rm --cpuset-cpus=0-1 app:latest python -m app.selfcheck
# Memory limit without swap, so the OOM killer behaves as in a pod
docker run --rm --memory=256m --memory-swap=256m app:latest python -m app.loadtest
For Lambda, the AWS base images include the Runtime Interface Emulator: running the function image locally exposes the invocation API on a port, so two invocations in a row — the minimum test for loop-binding mistakes — take a few seconds. docker pause and docker unpause imitate the freeze between invocations, which is how the stale-connection behaviour in this section was measured. A self-check entry point that prints the effective CPU and memory figures, as the integrated example below logs at startup, turns most configuration mistakes into a single line of evidence.
Resource boundaries¶
- Grace period: the drain must finish inside it — 10 s by default for
docker stop, 30 s for Kubernetes pods — with margin for the platform's own steps. - CPU: one server worker per allowed CPU as a starting point; process pools no larger than the limit; latency-critical services throttled in under a few percent of periods.
- Memory: every work-holding structure bounded from
memory.maxwith 25–30% headroom; a runtime guard at about 80%. - Lambda pools: sized for one invocation's concurrency, remembering each concurrent environment holds its own.
- Lambda invocation: every task finished before the handler returns.
- Startup visibility: log reported versus effective CPUs and the memory limit on every start, so a missing limit or a wrong variable is visible in the first line of output rather than in latency graphs days later.
Kubernetes specifics¶
Kubernetes layers its own lifecycle on top of these mechanisms, and two details interact with asyncio services in particular. First, termination is not instantaneous from the network's point of view: a pod receives SIGTERM at roughly the same time as it is removed from service endpoints, and load balancers may keep sending it traffic for a few seconds. A preStop hook that sleeps briefly, or a handler that keeps serving for a moment before draining, avoids rejecting those late requests — the sequence is in shutting down asyncio pods in Kubernetes. Second, probes run against the same event loop as requests: a liveness probe that times out because the loop is throttled or blocked will restart a pod that was merely slow, turning a capacity problem into a crash loop. Keep liveness probes cheap and generous, and put load-sensitive checks in readiness probes, as described in implementing health and readiness probes for asyncio.
Resource settings complete the picture: requests drive scheduling, limits drive throttling and OOM kills. For memory, requests close to limits; for CPU, requests from average use and limits with room for bursts — or no CPU limit where the cluster's policy allows, relying on requests for fair sharing. The downward API can pass both limits into the container as environment variables, which saves the application from parsing cgroup files.
Integrated production example¶
A small runtime module that a service imports at startup: it derives CPU and memory budgets from the cgroup, sizes the default executor and the work queue from them, installs the SIGTERM handler, and monitors throttling and memory pressure. Run in a container limited to two CPUs and 512 MiB on the 24-core host, it logged host cpu_count=24 -> effective cpus=2; memory limit=512 MiB; default executor=12; max queued=119, and docker stop completed in 0.22 s with exit code 0 and its cleanup logged:
import asyncio
import logging
import math
import os
import signal
from concurrent.futures import ThreadPoolExecutor
log = logging.getLogger("runtime")
def _cg(name: str) -> str | None:
try:
with open(f"/sys/fs/cgroup/{name}") as f:
return f.read().strip()
except OSError:
return None
def effective_cpus() -> int:
cpus = [len(os.sched_getaffinity(0))]
if (raw := _cg("cpu.max")) and not raw.startswith("max"):
quota, period = map(int, raw.split())
cpus.append(math.ceil(quota / period))
return max(1, min(cpus))
def memory_limit() -> int | None:
raw = _cg("memory.max")
return int(raw) if raw and raw != "max" else None
class ContainerRuntime:
"""Sizes and guards derived from the container's limits, not the host's."""
def __init__(self, per_item_bytes: int, baseline_bytes: int, headroom: float = 0.7) -> None:
self.cpus = effective_cpus()
self.mem_limit = memory_limit() or 1 << 30
self.max_queued = max(16, int((self.mem_limit * headroom - baseline_bytes) / per_item_bytes))
self.executor = ThreadPoolExecutor(max_workers=self.cpus * 4 + 4, thread_name_prefix="default")
self.over_memory = False
def install(self, loop: asyncio.AbstractEventLoop, main_task: asyncio.Task) -> None:
loop.set_default_executor(self.executor)
loop.add_signal_handler(signal.SIGTERM, main_task.cancel) # PID 1 needs this
log.info("host cpu_count=%s -> effective cpus=%s; memory limit=%d MiB; "
"default executor=%d; max queued=%d", os.cpu_count(), self.cpus,
self.mem_limit >> 20, self.executor._max_workers, self.max_queued)
async def monitor(self, interval: float = 1.0) -> None:
prev = _stat()
while True:
await asyncio.sleep(interval)
used = int(_cg("memory.current") or 0)
self.over_memory = used > 0.8 * self.mem_limit # callers shed load
cur = _stat()
if delta := cur.get("nr_throttled", 0) - prev.get("nr_throttled", 0):
log.warning("CPU throttled in %d periods in the last %.0fs", delta, interval)
prev = cur
def _stat() -> dict[str, int]:
raw = _cg("cpu.stat") or ""
return {k: int(v) for k, v in (line.split() for line in raw.splitlines())}
The service's main creates the runtime first, installs it with its own task, creates its queue with maxsize=rt.max_queued, starts rt.monitor() as a background task, and checks rt.over_memory before accepting work. The executor's size (12 here, from two effective CPUs) replaces the 28 threads Python would have created from the host's core count. The queue bound comes from the same arithmetic as the memory guide: budget times headroom, minus the measured baseline, divided by the per-item cost.
Diagnostic hook callout¶
Four series per container explain most runtime-induced incidents:
- Exit codes and
OOMKilledper termination — 137 withOOMKilled=falseon routine stops means signals are not reaching a handler;OOMKilled=truemeans a bound is missing. - Throttled periods per minute from
cpu.stat, next to event-loop lag p99. memory.current / memory.maxnext to the shed-load rate.- Lambda: whether each invocation created a new session or connection, from a log field.
Alert on any OOM kill, on throttling in more than 5% of periods at normal traffic, on memory above 80% of the limit for more than a minute, and on rollouts whose terminations average close to the grace period.
Failure modes¶
| Failure mode | Root cause | Detection | Fix |
|---|---|---|---|
| Stops take the full grace period, exit 137 | PID 1 without a handler, or a shell as PID 1 | Stop time, exit codes | SIGTERM handler; exec form; init |
| Too many threads/workers | Host CPU count used for sizing | Thread count vs CPU limit | Effective CPUs from cgroup; PYTHON_CPU_COUNT |
| Latency spikes at moderate CPU | CFS quota throttling | cpu.stat vs loop lag |
Higher limit, CPU work elsewhere |
| Sudden restarts, lost in-flight work | OOM kill | OOMKilled=true |
Bounds from memory.max; guard |
Event loop is closed in Lambda |
Loop-bound object across asyncio.run |
Second invocation fails | Module-level loop |
| Errors after idle in Lambda | Pooled connection closed while frozen | First call after idle fails | Client reconnect; retry writes with idempotency keys |
Frequently Asked Questions¶
Why does my asyncio container take so long to stop?
Python is PID 1 without a SIGTERM handler (the kernel ignores the signal), or a shell is PID 1 and does not forward it. In testing that took 10.26 s and ended in SIGKILL; with a handler, 0.44 s and exit 0.
Does Python respect Docker CPU and memory limits?
Not by itself: inside a --cpus=2, 256 MiB container it reported 24 CPUs and 60.9 GiB. Read the cgroup files or set PYTHON_CPU_COUNT (3.13+) and size pools and queues from them.
Why does my async service have latency spikes under a CPU limit?
The quota stops all threads once the per-period budget is used; under 0.5 CPU, loop lag p99 rose to 59.6 ms in testing. Check nr_throttled in /sys/fs/cgroup/cpu.stat.
How should I run asyncio in AWS Lambda?
Either asyncio.run per invocation with all async resources created inside, or a module-level event loop with loop.run_until_complete to reuse sessions; never reuse loop-bound objects across asyncio.run calls.
What happens to async cleanup when a container is OOM-killed?
Nothing runs: SIGKILL ends the process immediately. Bound memory use from the cgroup limit and shed load before reaching it.
Related¶
- Handling signals as PID 1 in containers — fast, clean stops.
- Sizing asyncio for container CPU limits — the CPUs you actually have.
- Detecting CPU throttling in async services — quota stalls, measured.
- Respecting container memory limits in async services — staying under the OOM line.
- Running asyncio in AWS Lambda handlers — loops and invocations.
- Reusing connections across Lambda invocations — reuse that fails safe.
- Graceful Shutdown & Signal Handling — the shutdown sequence itself.
- Resilience, Cancellation & Error Handling — the parent section.