Respecting Container Memory Limits in Async Services¶
Async services fail by memory in a characteristic way: a queue, a buffer or a set of in-flight requests grows faster than it drains, and the process keeps accepting work because nothing in it knows how much memory it is allowed. In a container, the kernel enforces that limit with the out-of-memory killer, which sends SIGKILL: no exception, no finally, no log line from the application. Measured with Docker on Linux, python:3.12-slim, in a container limited to 256 MiB: Python's view of system memory (os.sysconf) reported 60.9 GiB — the host's; /sys/fs/cgroup/memory.max reported the real 268,435,456 bytes. A program that queued 8 MiB items without consuming them was OOM-killed after 0.9 s, exit code 137, OOMKilled=true, and its finally block never ran. The same program checking memory.current against 80% of memory.max before each item stopped at 212 MiB after 24 items, ran its finally, and exited with status 2. This guide makes async services aware of their memory budget.
Prerequisites¶
- Docker or Kubernetes with cgroup v2; Python 3.11+.
- Bounded queues, from bounded asyncio.Queue with backpressure under load.
- The topic overview, Containers & Serverless.
1. Read the limit from the cgroup, not the OS¶
Every in-process way of asking "how much memory is there" returns the host's figure inside a container. The container's own limit and usage are in its cgroup:
import os
def cgroup_int(name: str) -> int | None:
try:
value = open(f"/sys/fs/cgroup/{name}").read().strip()
except OSError:
return None
return None if value == "max" else int(value)
host_total = os.sysconf("SC_PAGE_SIZE") * os.sysconf("SC_PHYS_PAGES") # 60.9 GiB: the host
limit = cgroup_int("memory.max") # 268435456: 256 MiB
used = cgroup_int("memory.current") # this container's usage
Measured inside a --memory=256m container: 60.9 GiB from sysconf, 256 MiB from memory.max. Libraries that size caches or buffers as a fraction of "system memory" therefore size them for the host, and so does any code that uses psutil.virtual_memory(). memory.current includes the page cache charged to the container as well as process memory, which is what the OOM killer acts on; memory.stat breaks it down if needed. On cgroup v1 hosts the equivalent files are memory.limit_in_bytes and memory.usage_in_bytes.
Verify: the service logs its cgroup memory limit at startup, and it matches the deployment's configured limit.
2. Understand what an OOM kill does to async cleanup¶
The kill is SIGKILL to the process; nothing in it runs afterwards:
async def main() -> None:
queue: asyncio.Queue = asyncio.Queue() # unbounded
try:
while True:
queue.put_nowait(await next_message()) # producer faster than consumers
finally:
await flush_and_close() # never runs after an OOM kill
Measured: exit code 137, OOMKilled=true in docker inspect, and no output from the finally. Everything the graceful-shutdown work in this section protects — draining requests, committing offsets, flushing telemetry — is skipped, and every in-flight request fails at once. Worse, the platform restarts the container and the same backlog can kill it again. The only defence is not reaching the limit: bound what can grow, and react before the kernel does. The growth patterns themselves are catalogued in diagnosing unbounded queue memory growth.
Verify: a staging load test at twice normal traffic produces backpressure or rejections, not exit code 137.
3. Bound queues and concurrency from the budget¶
Size every structure that holds work in memory from the container's limit and the size of a unit of work:
LIMIT = cgroup_int("memory.max") or 1 << 30 # fall back to 1 GiB outside containers
BASELINE = 120 * 2**20 # measured idle RSS of the service
PER_ITEM = 2 * 2**20 # measured memory per queued request
HEADROOM = 0.7
MAX_QUEUED = max(16, int((LIMIT * HEADROOM - BASELINE) / PER_ITEM))
queue: asyncio.Queue = asyncio.Queue(maxsize=MAX_QUEUED) # 298 for a 1 GiB container
A bounded queue converts excess load into backpressure — producers wait, or requests are rejected — instead of memory growth. The same arithmetic applies to the number of concurrent requests a server accepts, the number of streaming responses held open, and the size of in-process caches. Measure PER_ITEM under realistic payloads rather than estimating it, and keep headroom for the interpreter, libraries and fragmentation. Bounding concurrency in servers is covered in limiting concurrent requests with asyncio.Semaphore.
Verify: at the configured maximums, measured peak memory stays below 70–80% of the limit.
4. Shed load before the kernel does¶
Static bounds assume average payloads; a burst of unusually large ones can still approach the limit. A runtime guard that samples memory.current lets the service refuse work instead of dying:
class MemoryGuard:
def __init__(self, soft_fraction: float = 0.8) -> None:
self.limit = cgroup_int("memory.max")
self.soft = soft_fraction
self.over = False
async def run(self, interval: float = 0.5) -> None:
while True:
used = cgroup_int("memory.current")
if self.limit and used is not None:
self.over = used > self.soft * self.limit
MEMORY_FRACTION.set(used / self.limit)
await asyncio.sleep(interval)
guard = MemoryGuard()
async def accept(request):
if guard.over:
return Response(status=503, headers={"Retry-After": "2"}) # shed, do not die
return await handle(request)
Measured in the queue test: with the check at 80% of memory.max, the program stopped at 212 MiB, ran its finally, and exited with a status it chose. In a server, the equivalent is returning 503 with Retry-After while over the threshold, which spreads the load to other replicas and lets in-flight work finish. Sampling twice a second is cheap — a file read — and fast enough for most growth rates; the fastest-growing structures should still have static bounds.
Verify: under a deliberate memory spike, the service returns 503s and recovers, and its pod's restart count does not increase.
5. Set requests and limits from measurement¶
The platform's settings complete the picture. Measure peak memory under realistic load and set the limit with headroom above it:
resources:
requests:
memory: "768Mi" # what the scheduler reserves: near measured peak
limits:
memory: "1Gi" # the OOM line: peak plus headroom
env:
- name: MEMORY_LIMIT_BYTES
valueFrom: {resourceFieldRef: {resource: limits.memory}}
Unlike CPU, memory cannot be throttled — exceeding the limit means the OOM killer — so memory requests and limits are usually set close together, with the limit above measured peak. Passing the limit through the downward API gives the service its budget even where reading cgroup files is inconvenient. Track OOM kills as a first-class alert; each one means in-flight work was lost without any of the graceful paths running. The CPU counterpart is in sizing asyncio for container CPU limits.
Verify: OOM kills are alerted on, and the service's own memory-fraction gauge shows peaks well below the limit in normal operation.
Verification¶
An async service respects its memory limit when:
- It reads its limit from the cgroup (or the downward API), not from system memory.
- Every structure holding work is bounded from the budget and measured per-item costs.
- A runtime guard sheds load before usage reaches the limit.
- Limits come from measured peaks, and OOM kills are alerted on.
Diagnostic Hook: chart memory.current / memory.max per replica with the OOM-kill count and the 503 rate. A rising fraction followed by 503s is the guard working; a rising fraction followed by a restart is a bound that is missing or set too high.
Pitfalls & edge cases¶
- Trusting system memory figures. Measured: 60.9 GiB reported in a 256 MiB container.
- Unbounded queues. Measured: OOM-killed in 0.9 s, exit 137, cleanup skipped.
- Guards without static bounds. Fast growth can outrun a 0.5 s sampling interval.
- Ignoring the page cache.
memory.currentincludes it, and so does the OOM decision.
Frequently Asked Questions¶
Does Python know the Docker memory limit?
No. os.sysconf and psutil report the host's memory — 60.9 GiB inside a 256 MiB container in testing. Read /sys/fs/cgroup/memory.max (cgroup v2) or pass the limit in an environment variable.
What happens to asyncio cleanup when a container is OOM-killed?
Nothing runs: the kernel sends SIGKILL. In testing the container exited with 137 and OOMKilled=true, and the program's finally block never executed.
How do I prevent OOM kills in an async Python service?
Bound every queue and concurrency limit from the container's memory budget, and add a runtime guard that sheds load when memory.current passes about 80% of memory.max; in testing the guard stopped at 212 MiB and exited cleanly.
Should memory requests equal limits in Kubernetes?
Usually close to each other: memory cannot be throttled, so set the limit from measured peak plus headroom, and the request near the measured peak.
Related¶
- Containers & Serverless — up to the topic overview.
- Handling signals as PID 1 in containers — the graceful path an OOM kill skips.
- Resilience, Cancellation & Error Handling — the section overview.