Skip to content

Shutting Down asyncio Pods in Kubernetes

When Kubernetes stops a pod, two things happen in parallel: the kubelet sends SIGTERM to the container, and the endpoint controller removes the pod from Services so load balancers stop routing to it. Nothing orders them. A service that closes its listening socket the moment SIGTERM arrives refuses the requests that are still being routed to it during the seconds the endpoint removal takes to propagate. Replayed locally with an aiohttp service receiving a request every 50 ms, each taking 2 s, and a client that kept sending for 2.5 s after SIGTERM — the window in which a real load balancer may not yet know: closing immediately on SIGTERM let the 38 in-flight requests finish but refused 52 of 90 in total with connection errors; failing the readiness probe, waiting 3 s, and only then closing served all 90. This guide sets up the shutdown sequence, the probe and the pod settings that make rolling deploys invisible to clients.

Prerequisites

1. See the race between SIGTERM and endpoint removal

On pod deletion, the kubelet runs any preStop hook and then sends SIGTERM; independently, the endpoint controller updates Endpoints/EndpointSlices, and every kube-proxy, ingress controller and service mesh sidecar then updates its own routing. That propagation commonly takes one to several seconds:

async def main() -> None:
    runner = web.AppRunner(app, shutdown_timeout=30)
    await runner.setup()
    await web.TCPSite(runner, "0.0.0.0", 8080).start()

    stop = asyncio.Event()
    asyncio.get_running_loop().add_signal_handler(signal.SIGTERM, stop.set)
    await stop.wait()
    await runner.cleanup()           # naive: stop accepting immediately

Measured with this naive shutdown: the 38 requests already in flight completed — aiohttp's cleanup waits for them — but every request sent after SIGTERM was refused, 52 of 90 overall. In a real cluster the refused requests show up as 502s or connection resets at the ingress during every rolling deploy, usually blamed on the network.

Verify: run a load test through the ingress during a rolling restart; any 502s or resets at the moment pods terminate are this race.

Requests served during a pod shutdown, 90 sent 3 horizontal bars comparing close on SIGTERM: served with the others. Requests served during a pod shutdown, 90 sent close on SIGTERM: served 38 close on SIGTERM: refused 52 refused fail readiness, wait 3 s, close: served 90 aiohttp on Python 3.14; a request every 50 ms, each taking 2 s; client kept sending for 2.5 s after SIGTERM. The fix is to keep accepting until traffic has stopped arriving.

2. Fail readiness first, then wait

On SIGTERM, keep serving but report not-ready, so the endpoint is removed while the server still accepts connections. Then wait long enough for removal to propagate before closing:

ready = True


async def readyz(request: web.Request) -> web.Response:
    return web.Response(status=200 if ready else 503)


async def main() -> None:
    global ready
    ...
    await stop.wait()
    ready = False                          # 1. readiness probe now fails
    await asyncio.sleep(DRAIN_DELAY)       # 2. keep serving while routing updates (e.g. 5 s)
    await runner.cleanup()                 # 3. stop accepting; finish in-flight requests
    await close_resources()                # 4. pools, clients, telemetry

Measured with a 3 s delay: all 90 requests succeeded, including the 33 still in flight when the server finally closed. Kubernetes removes an endpoint either when the pod is marked terminating or when readiness fails, whichever the controller sees first; the delay covers propagation to every proxy. Choose DRAIN_DELAY from observed propagation in your cluster — 5 to 10 seconds is common — and longer behind external load balancers with their own health checks.

Verify: ingress logs during a rolling restart show the terminating pod's request rate falling to zero before its server closes.

3. Use a preStop hook when the app cannot delay itself

If the process cannot easily delay its own shutdown — a third-party server that exits on SIGTERM, or several containers in the pod — a preStop hook delays SIGTERM itself:

spec:
  terminationGracePeriodSeconds: 45
  containers:
    - name: api
      lifecycle:
        preStop:
          sleep:
            seconds: 5            # Kubernetes 1.30+: built-in sleep action
          # older clusters: exec: {command: ["sleep", "5"]} (needs a sleep binary in the image)
      readinessProbe:
        httpGet: {path: /readyz, port: 8080}
        periodSeconds: 2
        failureThreshold: 1

The kubelet runs preStop before sending SIGTERM, and the endpoint removal starts at the same moment as the hook — so the sleep gives routing time to update while the app keeps serving normally. The grace period must cover the preStop sleep, the in-flight drain and resource cleanup; when it expires, the kubelet sends SIGKILL. Uvicorn, for example, then only needs --timeout-graceful-shutdown for the drain itself.

Verify: kubectl describe pod during termination shows the preStop hook completing before the container receives SIGTERM, and the total stays under terminationGracePeriodSeconds.

A clean pod termination 2 lanes over time. A clean pod termination Kubernetes endpoint removal propagates no new traffic SIGKILL pod preStop sleep / not ready drain in-flight close pools exited time from pod deletion (grace period 45 s) → Keep serving until routing has caught up; only then stop accepting.

4. Drain in-flight work, then close resources in order

After routing has stopped, the remaining work is what is already in flight. Let it finish, bounded by what remains of the grace period:

async def shutdown(runner, background: set[asyncio.Task], budget: float) -> None:
    deadline = asyncio.get_running_loop().time() + budget
    await runner.shutdown()                                   # stop accepting new connections
    await runner.cleanup()                                    # waits for handlers (shutdown_timeout)
    for task in background:
        task.cancel("shutdown")
    remaining = max(0.0, deadline - asyncio.get_running_loop().time())
    await asyncio.wait(background, timeout=remaining)
    await db_pool.close()
    await http_client.aclose()
    await flush_telemetry()

Order matters: stop new work, finish or cancel in-flight work, then close the pools that work was using, then flush logs and metrics last so they capture the shutdown itself. Long-lived connections need explicit handling — WebSockets are closed with a "going away" code so clients reconnect elsewhere, as in closing WebSocket connections on shutdown, and the details of request draining are in draining in-flight requests before shutdown.

Verify: shutdown logs show each step with timestamps, finishing before the grace period.

5. Budget the grace period

terminationGracePeriodSeconds (default 30) is the hard limit for everything after the pod is deleted. Budget it explicitly:

# grace period = preStop/drain delay + longest request + background cancellation + cleanup + margin
DRAIN_DELAY = 5          # routing propagation
LONGEST_REQUEST = 20     # p99.9 of request duration, or the request timeout
BACKGROUND_STOP = 5      # cancel and await workers
CLEANUP = 3              # pools, clients, telemetry flush
MARGIN = 5
GRACE = DRAIN_DELAY + LONGEST_REQUEST + BACKGROUND_STOP + CLEANUP + MARGIN     # 38 -> set 45

Requests that can run longer than the budget — exports, uploads — should either move to background jobs that survive the pod, or be interrupted deliberately with a response the client can retry. Measure actual shutdown durations in production; if they approach the grace period, SIGKILLs will start cutting off cleanup, which shows up as lost telemetry and connections left half-open on the database side.

Verify: the 99th percentile of measured shutdown duration is well below terminationGracePeriodSeconds.

Which shutdown mechanism does this pod need? A decision on How does the app behave on SIGTERM with 4 outcomes. Which shutdown mechanism does this pod need? How does the app behave on SIGTERM? app controls shutdown readiness 503, delay, close in-process server exits at once preStop sleep hook delays SIGTERM long requests / WebSockets grace = delay + longest + cleanup budgeted work longer than grace background jobs not tied to the pod Every variant keeps accepting until the load balancer has moved on.

Verification

Pod shutdown is clean when:

  • Readiness fails or preStop delays before the server stops accepting.
  • The delay covers endpoint propagation, measured in your cluster.
  • In-flight work drains, then resources close in order, within the budget.
  • The grace period is sized from delay, longest request and cleanup.

Diagnostic Hook: during rolling deploys, chart ingress 5xx and connection errors per terminating pod, and log the timestamps of SIGTERM, readiness flip, server close and process exit. Errors clustered right after SIGTERM mean the server closes before routing updates; processes that never log their exit were SIGKILLed — the grace period is too short.

Pitfalls & edge cases

  • Closing the listener on SIGTERM. Measured: 52 of 90 requests refused.
  • Delay shorter than propagation. A shorter refusal window, but still errors.
  • Grace period shorter than drain plus cleanup. SIGKILL cuts off telemetry and connections.
  • preStop exec hooks without a sleep binary. Distroless images need the built-in sleep action.

Frequently Asked Questions

Why do I get 502 errors during Kubernetes rolling deploys?

Pods stop accepting connections on SIGTERM while load balancers still route to them, because endpoint removal propagates asynchronously. In a replay, closing on SIGTERM refused 52 of 90 requests; delaying the close served all 90.

How long should a preStop sleep be?

Long enough for endpoint removal to reach every proxy and load balancer in your cluster, often 5 to 10 seconds; measure it during a rolling deploy.

Should an asyncio service fail its readiness probe on SIGTERM?

Yes. Return 503 from readiness, keep serving for the drain delay, then stop accepting and finish in-flight requests.

What should terminationGracePeriodSeconds be?

At least the drain delay plus the longest request you will wait for, plus time to stop background work and close resources, with a margin. The default is 30 seconds.