Shutting Down asyncio Pods in Kubernetes¶
When Kubernetes stops a pod, two things happen in parallel: the kubelet sends SIGTERM to the container, and the endpoint controller removes the pod from Services so load balancers stop routing to it. Nothing orders them. A service that closes its listening socket the moment SIGTERM arrives refuses the requests that are still being routed to it during the seconds the endpoint removal takes to propagate. Replayed locally with an aiohttp service receiving a request every 50 ms, each taking 2 s, and a client that kept sending for 2.5 s after SIGTERM — the window in which a real load balancer may not yet know: closing immediately on SIGTERM let the 38 in-flight requests finish but refused 52 of 90 in total with connection errors; failing the readiness probe, waiting 3 s, and only then closing served all 90. This guide sets up the shutdown sequence, the probe and the pod settings that make rolling deploys invisible to clients.
Prerequisites¶
- Python 3.11+, an asyncio HTTP server (aiohttp, Uvicorn) running in Kubernetes.
- SIGTERM handling, from handling SIGTERM in asyncio services.
- Probes, from implementing health and readiness probes for asyncio.
1. See the race between SIGTERM and endpoint removal¶
On pod deletion, the kubelet runs any preStop hook and then sends SIGTERM; independently, the endpoint controller updates Endpoints/EndpointSlices, and every kube-proxy, ingress controller and service mesh sidecar then updates its own routing. That propagation commonly takes one to several seconds:
async def main() -> None:
runner = web.AppRunner(app, shutdown_timeout=30)
await runner.setup()
await web.TCPSite(runner, "0.0.0.0", 8080).start()
stop = asyncio.Event()
asyncio.get_running_loop().add_signal_handler(signal.SIGTERM, stop.set)
await stop.wait()
await runner.cleanup() # naive: stop accepting immediately
Measured with this naive shutdown: the 38 requests already in flight completed — aiohttp's cleanup waits for them — but every request sent after SIGTERM was refused, 52 of 90 overall. In a real cluster the refused requests show up as 502s or connection resets at the ingress during every rolling deploy, usually blamed on the network.
Verify: run a load test through the ingress during a rolling restart; any 502s or resets at the moment pods terminate are this race.
2. Fail readiness first, then wait¶
On SIGTERM, keep serving but report not-ready, so the endpoint is removed while the server still accepts connections. Then wait long enough for removal to propagate before closing:
ready = True
async def readyz(request: web.Request) -> web.Response:
return web.Response(status=200 if ready else 503)
async def main() -> None:
global ready
...
await stop.wait()
ready = False # 1. readiness probe now fails
await asyncio.sleep(DRAIN_DELAY) # 2. keep serving while routing updates (e.g. 5 s)
await runner.cleanup() # 3. stop accepting; finish in-flight requests
await close_resources() # 4. pools, clients, telemetry
Measured with a 3 s delay: all 90 requests succeeded, including the 33 still in flight when the server finally closed. Kubernetes removes an endpoint either when the pod is marked terminating or when readiness fails, whichever the controller sees first; the delay covers propagation to every proxy. Choose DRAIN_DELAY from observed propagation in your cluster — 5 to 10 seconds is common — and longer behind external load balancers with their own health checks.
Verify: ingress logs during a rolling restart show the terminating pod's request rate falling to zero before its server closes.
3. Use a preStop hook when the app cannot delay itself¶
If the process cannot easily delay its own shutdown — a third-party server that exits on SIGTERM, or several containers in the pod — a preStop hook delays SIGTERM itself:
spec:
terminationGracePeriodSeconds: 45
containers:
- name: api
lifecycle:
preStop:
sleep:
seconds: 5 # Kubernetes 1.30+: built-in sleep action
# older clusters: exec: {command: ["sleep", "5"]} (needs a sleep binary in the image)
readinessProbe:
httpGet: {path: /readyz, port: 8080}
periodSeconds: 2
failureThreshold: 1
The kubelet runs preStop before sending SIGTERM, and the endpoint removal starts at the same moment as the hook — so the sleep gives routing time to update while the app keeps serving normally. The grace period must cover the preStop sleep, the in-flight drain and resource cleanup; when it expires, the kubelet sends SIGKILL. Uvicorn, for example, then only needs --timeout-graceful-shutdown for the drain itself.
Verify: kubectl describe pod during termination shows the preStop hook completing before the container receives SIGTERM, and the total stays under terminationGracePeriodSeconds.
4. Drain in-flight work, then close resources in order¶
After routing has stopped, the remaining work is what is already in flight. Let it finish, bounded by what remains of the grace period:
async def shutdown(runner, background: set[asyncio.Task], budget: float) -> None:
deadline = asyncio.get_running_loop().time() + budget
await runner.shutdown() # stop accepting new connections
await runner.cleanup() # waits for handlers (shutdown_timeout)
for task in background:
task.cancel("shutdown")
remaining = max(0.0, deadline - asyncio.get_running_loop().time())
await asyncio.wait(background, timeout=remaining)
await db_pool.close()
await http_client.aclose()
await flush_telemetry()
Order matters: stop new work, finish or cancel in-flight work, then close the pools that work was using, then flush logs and metrics last so they capture the shutdown itself. Long-lived connections need explicit handling — WebSockets are closed with a "going away" code so clients reconnect elsewhere, as in closing WebSocket connections on shutdown, and the details of request draining are in draining in-flight requests before shutdown.
Verify: shutdown logs show each step with timestamps, finishing before the grace period.
5. Budget the grace period¶
terminationGracePeriodSeconds (default 30) is the hard limit for everything after the pod is deleted. Budget it explicitly:
# grace period = preStop/drain delay + longest request + background cancellation + cleanup + margin
DRAIN_DELAY = 5 # routing propagation
LONGEST_REQUEST = 20 # p99.9 of request duration, or the request timeout
BACKGROUND_STOP = 5 # cancel and await workers
CLEANUP = 3 # pools, clients, telemetry flush
MARGIN = 5
GRACE = DRAIN_DELAY + LONGEST_REQUEST + BACKGROUND_STOP + CLEANUP + MARGIN # 38 -> set 45
Requests that can run longer than the budget — exports, uploads — should either move to background jobs that survive the pod, or be interrupted deliberately with a response the client can retry. Measure actual shutdown durations in production; if they approach the grace period, SIGKILLs will start cutting off cleanup, which shows up as lost telemetry and connections left half-open on the database side.
Verify: the 99th percentile of measured shutdown duration is well below terminationGracePeriodSeconds.
Verification¶
Pod shutdown is clean when:
- Readiness fails or
preStopdelays before the server stops accepting. - The delay covers endpoint propagation, measured in your cluster.
- In-flight work drains, then resources close in order, within the budget.
- The grace period is sized from delay, longest request and cleanup.
Diagnostic Hook: during rolling deploys, chart ingress 5xx and connection errors per terminating pod, and log the timestamps of SIGTERM, readiness flip, server close and process exit. Errors clustered right after SIGTERM mean the server closes before routing updates; processes that never log their exit were SIGKILLed — the grace period is too short.
Pitfalls & edge cases¶
- Closing the listener on SIGTERM. Measured: 52 of 90 requests refused.
- Delay shorter than propagation. A shorter refusal window, but still errors.
- Grace period shorter than drain plus cleanup. SIGKILL cuts off telemetry and connections.
- preStop exec hooks without a sleep binary. Distroless images need the built-in sleep action.
Frequently Asked Questions¶
Why do I get 502 errors during Kubernetes rolling deploys?
Pods stop accepting connections on SIGTERM while load balancers still route to them, because endpoint removal propagates asynchronously. In a replay, closing on SIGTERM refused 52 of 90 requests; delaying the close served all 90.
How long should a preStop sleep be?
Long enough for endpoint removal to reach every proxy and load balancer in your cluster, often 5 to 10 seconds; measure it during a rolling deploy.
Should an asyncio service fail its readiness probe on SIGTERM?
Yes. Return 503 from readiness, keep serving for the drain delay, then stop accepting and finish in-flight requests.
What should terminationGracePeriodSeconds be?
At least the drain delay plus the longest request you will wait for, plus time to stop background work and close resources, with a margin. The default is 30 seconds.
Related¶
- Graceful Shutdown & Signals — up to the topic overview.
- Flushing telemetry and logs before exit — the last step of the sequence.
- Resilience, Cancellation & Error Handling — the section overview.