Skip to content

Async in Containers & Serverless

Most asyncio services run inside a container or a serverless runtime, and both environments change rules that Python programs normally take for granted. The process may be PID 1, where unhandled signals are ignored. The CPU count and memory size the interpreter reports describe the host, not the allowance. CPU limits are enforced by stopping the process for parts of every 100 ms, and memory limits by killing it. In serverless, the event loop has to be bridged into a synchronous handler and survive freezes between invocations. This section measures each of these on Docker and the AWS Lambda Python base image. An asyncio program as PID 1 with no SIGTERM handler took 10.26 s to stop and was killed with exit 137; with a handler, 0.44 s and exit 0. In a --cpus=2 container, Python reported 24 CPUs and sized its default executor at 28 threads. Under a 0.5-CPU quota, a busy background thread pushed event-loop lag p99 from 10.6 to 59.6 ms, with 2.6 s of throttling in five seconds. In a 256 MiB container Python reported 60.9 GiB of memory, and an unbounded queue was OOM-killed in 0.9 s with no cleanup. In Lambda, asyncio.run per invocation took 0.87 ms warm against 0.19 ms with a module-level loop — and a global session reused across asyncio.run calls failed 20 of 20 invocations.

The parent section, Resilience, Cancellation & Error Handling, covers shutdown, timeouts and load shedding in general; this topic covers how container and serverless runtimes change them, and how to tell from inside a running service which of their rules is causing a problem.

Architectural principles

  • Handle SIGTERM explicitly, and make sure Python is the process that receives it.
  • Derive every size from the container's limits — CPU quota, affinity, memory — never from host-reported figures.
  • Treat throttling as a latency source and measure it directly from the cgroup.
  • Stay under the memory limit by construction, because the OOM killer runs no cleanup.
  • In serverless, the loop's lifetime decides what can be reused, and nothing may run between invocations.
What the runtime changes, layer by layer 4 stacked layers. What the runtime changes, layer by layer process start PID 1 ignores unhandled SIGTERM; sh hides it CPU quota stops all threads per 100 ms period; host core count reported memory limit enforced by SIGKILL; host memory reported serverless invocation sync handler, frozen between calls Each layer has one guide in this section.

Execution model: the runtime around the event loop

Inside a container, the event loop runs exactly as it does anywhere else; what differs is the environment's response to it. Signals: Docker and Kubernetes stop a container by signalling its PID 1. If Python is PID 1 and has no SIGTERM handler, the kernel ignores the signal — the PID-1 rule — and the platform kills the process after the grace period. If a shell is PID 1, the signal never reaches Python at all. CPU: a limit of 0.5 CPU means 50 ms of CPU time per 100 ms period for all of the process's threads together. A burst that uses the budget early stops every thread, including the one running the loop, until the period ends — invisible in average utilization, obvious in loop lag. Memory: the cgroup charges the container's usage, including page cache, against a hard limit; crossing it brings the OOM killer's SIGKILL.

Lambda adds an invocation model on top. The runtime calls a synchronous handler once per invocation; the execution environment, with its module-level state, is reused for later invocations and frozen between them. An event loop created by asyncio.run in the handler dies at the end of each invocation, taking with it anything bound to it; a loop created at module level lives as long as the environment. Nothing runs while frozen — not background tasks, not keep-alive timers — while servers on the other end of pooled connections keep counting their idle timeouts.

The measurements behind this section A grid of 6 rows by 3 columns. The measurements behind this section question measured guide stop without a handler? 10.26 s, exit 137; with handler 0.44 s, exit 0 PID 1 signals CPUs seen under --cpus=2? 24; executor 28 threads CPU limits lag under a 0.5-CPU quota? p99 59.6 ms; 51 throttled periods in 5 s throttling memory seen in 256 MiB? 60.9 GiB; OOM-killed in 0.9 s, no cleanup memory limits asyncio.run per Lambda call? 0.87 ms warm vs 0.19 ms module loop Lambda handlers pooled connection after a freeze? reconnected in 2.35 ms, no error Lambda connections Docker on Linux with python:3.12-slim; AWS Lambda Python 3.13 base image with its emulator.

Pattern catalogue

A SIGTERM handler, with Python as PID 1 (or behind an init)

loop.add_signal_handler(signal.SIGTERM, asyncio.current_task().cancel)

Exec-form CMD, exec at the end of entrypoint scripts, and tini when the application spawns children. The measurements separate the pieces: the handler alone fixed the stop (0.44 s, exit 0); --init alone made it fast but abrupt (0.22 s, exit 143, no cleanup); a shell-form command defeated even a correct handler (10.25 s, exit 137). See handling signals as PID 1 in containers.

Sizes from the cgroup

Effective CPUs from cpu.max and affinity (or PYTHON_CPU_COUNT on 3.13+), and pools and workers sized from them. Python 3.13's os.process_cpu_count() followed a cpuset (2 CPUs, 6 executor threads) but not a quota, which is how Kubernetes limits work, so the cgroup file or the environment variable is the reliable source. See sizing asyncio for container CPU limits.

Throttling as a metric

nr_throttled and throttled_usec from cpu.stat, plotted against loop lag; CPU work moved out of latency-critical processes. The same busy thread produced 10.6 ms of lag p99 without a limit (GIL contention, zero throttling) and 59.6 ms under 0.5 CPU (51 throttled periods in five seconds), so the counters, not the lag alone, identify the cause. See detecting CPU throttling in async services.

Memory bounds and a guard

Queues and concurrency sized from memory.max, and a guard that sheds load at 80% of it. The guard turned a 0.9-second OOM kill with no cleanup into a controlled stop at 212 MiB with the finally block run and an exit status the program chose. See respecting container memory limits in async services.

A module-level loop in Lambda

LOOP = asyncio.new_event_loop()

def handler(event, context):
    return LOOP.run_until_complete(work(event))

The module-level loop is what lets a session and its pooled connections survive from one invocation to the next: 0.19 ms per warm invocation against 0.87 ms for asyncio.run with a fresh session. After a 4-second freeze past the upstream's 2-second keep-alive timeout, aiohttp replaced the dead connection transparently (2.35 ms) — safe for reads; writes should retry once with an idempotency key. See running asyncio in AWS Lambda handlers and reusing connections across Lambda invocations.

Reproducing the runtime locally

Every behaviour in this section can be reproduced on a laptop with Docker, which makes them testable before they appear in production. The flags map directly onto the production mechanisms:

# PID 1 and signals: time the stop and check the exit code
docker run -d --name t app:latest && sleep 3 && time docker stop -t 10 t
docker inspect -f '{{.State.ExitCode}} OOMKilled={{.State.OOMKilled}}' t

# CPU quota and cpuset, as Kubernetes limits and pinned nodes would apply them
docker run --rm --cpus=0.5 app:latest python -m app.selfcheck
docker run --rm --cpuset-cpus=0-1 app:latest python -m app.selfcheck

# Memory limit without swap, so the OOM killer behaves as in a pod
docker run --rm --memory=256m --memory-swap=256m app:latest python -m app.loadtest

For Lambda, the AWS base images include the Runtime Interface Emulator: running the function image locally exposes the invocation API on a port, so two invocations in a row — the minimum test for loop-binding mistakes — take a few seconds. docker pause and docker unpause imitate the freeze between invocations, which is how the stale-connection behaviour in this section was measured. A self-check entry point that prints the effective CPU and memory figures, as the integrated example below logs at startup, turns most configuration mistakes into a single line of evidence.

Resource boundaries

  • Grace period: the drain must finish inside it — 10 s by default for docker stop, 30 s for Kubernetes pods — with margin for the platform's own steps.
  • CPU: one server worker per allowed CPU as a starting point; process pools no larger than the limit; latency-critical services throttled in under a few percent of periods.
  • Memory: every work-holding structure bounded from memory.max with 25–30% headroom; a runtime guard at about 80%.
  • Lambda pools: sized for one invocation's concurrency, remembering each concurrent environment holds its own.
  • Lambda invocation: every task finished before the handler returns.
  • Startup visibility: log reported versus effective CPUs and the memory limit on every start, so a missing limit or a wrong variable is visible in the first line of output rather than in latency graphs days later.

Kubernetes specifics

Kubernetes layers its own lifecycle on top of these mechanisms, and two details interact with asyncio services in particular. First, termination is not instantaneous from the network's point of view: a pod receives SIGTERM at roughly the same time as it is removed from service endpoints, and load balancers may keep sending it traffic for a few seconds. A preStop hook that sleeps briefly, or a handler that keeps serving for a moment before draining, avoids rejecting those late requests — the sequence is in shutting down asyncio pods in Kubernetes. Second, probes run against the same event loop as requests: a liveness probe that times out because the loop is throttled or blocked will restart a pod that was merely slow, turning a capacity problem into a crash loop. Keep liveness probes cheap and generous, and put load-sensitive checks in readiness probes, as described in implementing health and readiness probes for asyncio.

Resource settings complete the picture: requests drive scheduling, limits drive throttling and OOM kills. For memory, requests close to limits; for CPU, requests from average use and limits with room for bursts — or no CPU limit where the cluster's policy allows, relying on requests for fair sharing. The downward API can pass both limits into the container as environment variables, which saves the application from parsing cgroup files.

Which Kubernetes setting addresses this symptom? A decision on What is the symptom with 4 outcomes. Which Kubernetes setting addresses this symptom? What is the symptom? slow rollouts, exit 137 SIGTERM handler, exec form 10.26 s -> 0.44 s errors right after SIGTERM preStop delay before draining endpoints update late restarts under load cheap liveness, real readiness slow is not dead latency spikes, throttling higher CPU limit or less CPU work cpu.stat Most container incidents trace back to one of these four settings.

Integrated production example

A small runtime module that a service imports at startup: it derives CPU and memory budgets from the cgroup, sizes the default executor and the work queue from them, installs the SIGTERM handler, and monitors throttling and memory pressure. Run in a container limited to two CPUs and 512 MiB on the 24-core host, it logged host cpu_count=24 -> effective cpus=2; memory limit=512 MiB; default executor=12; max queued=119, and docker stop completed in 0.22 s with exit code 0 and its cleanup logged:

import asyncio
import logging
import math
import os
import signal
from concurrent.futures import ThreadPoolExecutor

log = logging.getLogger("runtime")


def _cg(name: str) -> str | None:
    try:
        with open(f"/sys/fs/cgroup/{name}") as f:
            return f.read().strip()
    except OSError:
        return None


def effective_cpus() -> int:
    cpus = [len(os.sched_getaffinity(0))]
    if (raw := _cg("cpu.max")) and not raw.startswith("max"):
        quota, period = map(int, raw.split())
        cpus.append(math.ceil(quota / period))
    return max(1, min(cpus))


def memory_limit() -> int | None:
    raw = _cg("memory.max")
    return int(raw) if raw and raw != "max" else None


class ContainerRuntime:
    """Sizes and guards derived from the container's limits, not the host's."""

    def __init__(self, per_item_bytes: int, baseline_bytes: int, headroom: float = 0.7) -> None:
        self.cpus = effective_cpus()
        self.mem_limit = memory_limit() or 1 << 30
        self.max_queued = max(16, int((self.mem_limit * headroom - baseline_bytes) / per_item_bytes))
        self.executor = ThreadPoolExecutor(max_workers=self.cpus * 4 + 4, thread_name_prefix="default")
        self.over_memory = False

    def install(self, loop: asyncio.AbstractEventLoop, main_task: asyncio.Task) -> None:
        loop.set_default_executor(self.executor)
        loop.add_signal_handler(signal.SIGTERM, main_task.cancel)       # PID 1 needs this
        log.info("host cpu_count=%s -> effective cpus=%s; memory limit=%d MiB; "
                 "default executor=%d; max queued=%d", os.cpu_count(), self.cpus,
                 self.mem_limit >> 20, self.executor._max_workers, self.max_queued)

    async def monitor(self, interval: float = 1.0) -> None:
        prev = _stat()
        while True:
            await asyncio.sleep(interval)
            used = int(_cg("memory.current") or 0)
            self.over_memory = used > 0.8 * self.mem_limit              # callers shed load
            cur = _stat()
            if delta := cur.get("nr_throttled", 0) - prev.get("nr_throttled", 0):
                log.warning("CPU throttled in %d periods in the last %.0fs", delta, interval)
            prev = cur


def _stat() -> dict[str, int]:
    raw = _cg("cpu.stat") or ""
    return {k: int(v) for k, v in (line.split() for line in raw.splitlines())}

The service's main creates the runtime first, installs it with its own task, creates its queue with maxsize=rt.max_queued, starts rt.monitor() as a background task, and checks rt.over_memory before accepting work. The executor's size (12 here, from two effective CPUs) replaces the 28 threads Python would have created from the host's core count. The queue bound comes from the same arithmetic as the memory guide: budget times headroom, minus the measured baseline, divided by the per-item cost.

Diagnostic hook callout

Four series per container explain most runtime-induced incidents:

  • Exit codes and OOMKilled per termination — 137 with OOMKilled=false on routine stops means signals are not reaching a handler; OOMKilled=true means a bound is missing.
  • Throttled periods per minute from cpu.stat, next to event-loop lag p99.
  • memory.current / memory.max next to the shed-load rate.
  • Lambda: whether each invocation created a new session or connection, from a log field.

Alert on any OOM kill, on throttling in more than 5% of periods at normal traffic, on memory above 80% of the limit for more than a minute, and on rollouts whose terminations average close to the grace period.

Failure modes

Failure mode Root cause Detection Fix
Stops take the full grace period, exit 137 PID 1 without a handler, or a shell as PID 1 Stop time, exit codes SIGTERM handler; exec form; init
Too many threads/workers Host CPU count used for sizing Thread count vs CPU limit Effective CPUs from cgroup; PYTHON_CPU_COUNT
Latency spikes at moderate CPU CFS quota throttling cpu.stat vs loop lag Higher limit, CPU work elsewhere
Sudden restarts, lost in-flight work OOM kill OOMKilled=true Bounds from memory.max; guard
Event loop is closed in Lambda Loop-bound object across asyncio.run Second invocation fails Module-level loop
Errors after idle in Lambda Pooled connection closed while frozen First call after idle fails Client reconnect; retry writes with idempotency keys

Frequently Asked Questions

Why does my asyncio container take so long to stop?

Python is PID 1 without a SIGTERM handler (the kernel ignores the signal), or a shell is PID 1 and does not forward it. In testing that took 10.26 s and ended in SIGKILL; with a handler, 0.44 s and exit 0.

Does Python respect Docker CPU and memory limits?

Not by itself: inside a --cpus=2, 256 MiB container it reported 24 CPUs and 60.9 GiB. Read the cgroup files or set PYTHON_CPU_COUNT (3.13+) and size pools and queues from them.

Why does my async service have latency spikes under a CPU limit?

The quota stops all threads once the per-period budget is used; under 0.5 CPU, loop lag p99 rose to 59.6 ms in testing. Check nr_throttled in /sys/fs/cgroup/cpu.stat.

How should I run asyncio in AWS Lambda?

Either asyncio.run per invocation with all async resources created inside, or a module-level event loop with loop.run_until_complete to reuse sessions; never reuse loop-bound objects across asyncio.run calls.

What happens to async cleanup when a container is OOM-killed?

Nothing runs: SIGKILL ends the process immediately. Bound memory use from the cgroup limit and shed load before reaching it.