Skip to content

Respecting Container Memory Limits in Async Services

Async services fail by memory in a characteristic way: a queue, a buffer or a set of in-flight requests grows faster than it drains, and the process keeps accepting work because nothing in it knows how much memory it is allowed. In a container, the kernel enforces that limit with the out-of-memory killer, which sends SIGKILL: no exception, no finally, no log line from the application. Measured with Docker on Linux, python:3.12-slim, in a container limited to 256 MiB: Python's view of system memory (os.sysconf) reported 60.9 GiB — the host's; /sys/fs/cgroup/memory.max reported the real 268,435,456 bytes. A program that queued 8 MiB items without consuming them was OOM-killed after 0.9 s, exit code 137, OOMKilled=true, and its finally block never ran. The same program checking memory.current against 80% of memory.max before each item stopped at 212 MiB after 24 items, ran its finally, and exited with status 2. This guide makes async services aware of their memory budget.

Prerequisites

1. Read the limit from the cgroup, not the OS

Every in-process way of asking "how much memory is there" returns the host's figure inside a container. The container's own limit and usage are in its cgroup:

import os


def cgroup_int(name: str) -> int | None:
    try:
        value = open(f"/sys/fs/cgroup/{name}").read().strip()
    except OSError:
        return None
    return None if value == "max" else int(value)


host_total = os.sysconf("SC_PAGE_SIZE") * os.sysconf("SC_PHYS_PAGES")   # 60.9 GiB: the host
limit = cgroup_int("memory.max")                                         # 268435456: 256 MiB
used = cgroup_int("memory.current")                                      # this container's usage

Measured inside a --memory=256m container: 60.9 GiB from sysconf, 256 MiB from memory.max. Libraries that size caches or buffers as a fraction of "system memory" therefore size them for the host, and so does any code that uses psutil.virtual_memory(). memory.current includes the page cache charged to the container as well as process memory, which is what the OOM killer acts on; memory.stat breaks it down if needed. On cgroup v1 hosts the equivalent files are memory.limit_in_bytes and memory.usage_in_bytes.

Verify: the service logs its cgroup memory limit at startup, and it matches the deployment's configured limit.

A queue of 8 MiB items in a 256 MiB container A grid of 3 rows by 4 columns. A queue of 8 MiB items in a 256 MiB container variant stopped at how it ended cleanup ran what Python sees sysconf: 60.9 GiB cgroup memory.max: 256 MiB n/a unbounded queue OOM after 0.9 s SIGKILL, exit 137, OOMKilled=true no guard at 80% of memory.max 212 MiB, 24 items exit 2, chosen by the app yes python:3.12-slim with --memory=256m --memory-swap=256m on Docker, Linux.

2. Understand what an OOM kill does to async cleanup

The kill is SIGKILL to the process; nothing in it runs afterwards:

async def main() -> None:
    queue: asyncio.Queue = asyncio.Queue()           # unbounded
    try:
        while True:
            queue.put_nowait(await next_message())    # producer faster than consumers
    finally:
        await flush_and_close()                       # never runs after an OOM kill

Measured: exit code 137, OOMKilled=true in docker inspect, and no output from the finally. Everything the graceful-shutdown work in this section protects — draining requests, committing offsets, flushing telemetry — is skipped, and every in-flight request fails at once. Worse, the platform restarts the container and the same backlog can kill it again. The only defence is not reaching the limit: bound what can grow, and react before the kernel does. The growth patterns themselves are catalogued in diagnosing unbounded queue memory growth.

Verify: a staging load test at twice normal traffic produces backpressure or rejections, not exit code 137.

3. Bound queues and concurrency from the budget

Size every structure that holds work in memory from the container's limit and the size of a unit of work:

LIMIT = cgroup_int("memory.max") or 1 << 30                 # fall back to 1 GiB outside containers
BASELINE = 120 * 2**20                                       # measured idle RSS of the service
PER_ITEM = 2 * 2**20                                         # measured memory per queued request
HEADROOM = 0.7

MAX_QUEUED = max(16, int((LIMIT * HEADROOM - BASELINE) / PER_ITEM))
queue: asyncio.Queue = asyncio.Queue(maxsize=MAX_QUEUED)     # 298 for a 1 GiB container

A bounded queue converts excess load into backpressure — producers wait, or requests are rejected — instead of memory growth. The same arithmetic applies to the number of concurrent requests a server accepts, the number of streaming responses held open, and the size of in-process caches. Measure PER_ITEM under realistic payloads rather than estimating it, and keep headroom for the interpreter, libraries and fragmentation. Bounding concurrency in servers is covered in limiting concurrent requests with asyncio.Semaphore.

Verify: at the configured maximums, measured peak memory stays below 70–80% of the limit.

Turning the memory limit into bounds A flow of 5 stages. Turning the memory limit into bounds read memory.max 256 MiB, not 60.9 GiB budget - baseline measured idle use / per-item cost measured, with headroom queue maxsize, semaphores bounded work watch memory.current shed at 80% Static bounds from the budget, a dynamic guard for surprises.

4. Shed load before the kernel does

Static bounds assume average payloads; a burst of unusually large ones can still approach the limit. A runtime guard that samples memory.current lets the service refuse work instead of dying:

class MemoryGuard:
    def __init__(self, soft_fraction: float = 0.8) -> None:
        self.limit = cgroup_int("memory.max")
        self.soft = soft_fraction
        self.over = False

    async def run(self, interval: float = 0.5) -> None:
        while True:
            used = cgroup_int("memory.current")
            if self.limit and used is not None:
                self.over = used > self.soft * self.limit
                MEMORY_FRACTION.set(used / self.limit)
            await asyncio.sleep(interval)


guard = MemoryGuard()


async def accept(request):
    if guard.over:
        return Response(status=503, headers={"Retry-After": "2"})   # shed, do not die
    return await handle(request)

Measured in the queue test: with the check at 80% of memory.max, the program stopped at 212 MiB, ran its finally, and exited with a status it chose. In a server, the equivalent is returning 503 with Retry-After while over the threshold, which spreads the load to other replicas and lets in-flight work finish. Sampling twice a second is cheap — a file read — and fast enough for most growth rates; the fastest-growing structures should still have static bounds.

Verify: under a deliberate memory spike, the service returns 503s and recovers, and its pod's restart count does not increase.

5. Set requests and limits from measurement

The platform's settings complete the picture. Measure peak memory under realistic load and set the limit with headroom above it:

resources:
  requests:
    memory: "768Mi"      # what the scheduler reserves: near measured peak
  limits:
    memory: "1Gi"        # the OOM line: peak plus headroom
env:
  - name: MEMORY_LIMIT_BYTES
    valueFrom: {resourceFieldRef: {resource: limits.memory}}

Unlike CPU, memory cannot be throttled — exceeding the limit means the OOM killer — so memory requests and limits are usually set close together, with the limit above measured peak. Passing the limit through the downward API gives the service its budget even where reading cgroup files is inconvenient. Track OOM kills as a first-class alert; each one means in-flight work was lost without any of the graceful paths running. The CPU counterpart is in sizing asyncio for container CPU limits.

Verify: OOM kills are alerted on, and the service's own memory-fraction gauge shows peaks well below the limit in normal operation.

How should this service stay inside its memory limit? A decision on What could grow with 4 outcomes. How should this service stay inside its memory limit? What could grow? anything sized from 'system memory' read memory.max instead 256 MiB, not 60.9 GiB queues, in-flight requests, caches static bounds from the budget measured per-item cost payload sizes vary widely runtime guard on memory.current shed at 80% the deployment itself limit = measured peak + headroom alert on OOM kills The OOM killer runs no cleanup; staying under the line is the only graceful path.

Verification

An async service respects its memory limit when:

  • It reads its limit from the cgroup (or the downward API), not from system memory.
  • Every structure holding work is bounded from the budget and measured per-item costs.
  • A runtime guard sheds load before usage reaches the limit.
  • Limits come from measured peaks, and OOM kills are alerted on.

Diagnostic Hook: chart memory.current / memory.max per replica with the OOM-kill count and the 503 rate. A rising fraction followed by 503s is the guard working; a rising fraction followed by a restart is a bound that is missing or set too high.

Pitfalls & edge cases

  • Trusting system memory figures. Measured: 60.9 GiB reported in a 256 MiB container.
  • Unbounded queues. Measured: OOM-killed in 0.9 s, exit 137, cleanup skipped.
  • Guards without static bounds. Fast growth can outrun a 0.5 s sampling interval.
  • Ignoring the page cache. memory.current includes it, and so does the OOM decision.

Frequently Asked Questions

Does Python know the Docker memory limit?

No. os.sysconf and psutil report the host's memory — 60.9 GiB inside a 256 MiB container in testing. Read /sys/fs/cgroup/memory.max (cgroup v2) or pass the limit in an environment variable.

What happens to asyncio cleanup when a container is OOM-killed?

Nothing runs: the kernel sends SIGKILL. In testing the container exited with 137 and OOMKilled=true, and the program's finally block never executed.

How do I prevent OOM kills in an async Python service?

Bound every queue and concurrency limit from the container's memory budget, and add a runtime guard that sheds load when memory.current passes about 80% of memory.max; in testing the guard stopped at 212 MiB and exited cleanly.

Should memory requests equal limits in Kubernetes?

Usually close to each other: memory cannot be throttled, so set the limit from measured peak plus headroom, and the request near the measured peak.