Detecting CPU Throttling in Async Services¶
A container CPU limit is not a slower CPU; it is a budget of CPU time per scheduling period — by default 100 ms. A process that spends its budget early in a period is stopped, all of its threads together, until the next period begins. For an asyncio service that means the event loop freezes for tens of milliseconds at a time, and every request in flight waits, even though average CPU usage looks comfortably under the limit. Measured with Docker on Linux, Python 3.12 in a container running an event-loop lag probe for five seconds, alongside a background thread doing CPU work — the shape of a to_thread job or a CPU-heavy library call: with no limit, loop lag p99 was 10.6 ms (the cost of sharing the GIL with a busy thread); with --cpus=2, also 10.6 ms; with --cpus=0.5, 59.6 ms p99 and 61.7 ms max, and the cgroup's cpu.stat recorded 51 throttled periods totalling 2,590 ms of throttling in five seconds. Without the background thread, lag stayed at 0.1–0.3 ms under every limit. This guide detects throttling from inside the service and removes its causes.
Prerequisites¶
- Docker or Kubernetes with cgroup v2; Python 3.11+.
- Loop lag as a metric, from measuring event loop lag in production.
- CPU limits and Python, from sizing asyncio for container CPU limits.
1. Read throttling counters from cpu.stat¶
The kernel counts throttling per cgroup. Inside a container, /sys/fs/cgroup/cpu.stat is the container's own:
def cpu_stat() -> dict[str, int]:
try:
with open("/sys/fs/cgroup/cpu.stat") as f:
return {k: int(v) for k, v in (line.split() for line in f)}
except OSError:
return {}
# fields include: usage_usec, nr_periods, nr_throttled, throttled_usec
before = cpu_stat()
await asyncio.sleep(5)
after = cpu_stat()
throttled_periods = after["nr_throttled"] - before["nr_throttled"]
throttled_ms = (after["throttled_usec"] - before["throttled_usec"]) / 1000
Measured over five seconds under --cpus=0.5 with the background thread: 51 throttled periods and 2,590 ms throttled — the process was stopped for more than half of the wall-clock time. Under --cpus=2 the same workload recorded zero, because one busy thread could not exhaust a two-CPU budget. nr_throttled / nr_periods is the fraction of periods in which the budget ran out; anything persistently above a few percent means requests are regularly being paused.
Verify: export the throttled-period rate and throttled time as metrics, sampled every few seconds.
2. Correlate throttling with loop lag¶
Throttling explains latency only if it coincides with it. Record both in the same sampling task, so a dashboard can put them side by side:
async def monitor(period: float = 5.0) -> None:
loop = asyncio.get_running_loop()
prev = cpu_stat()
while True:
lags = []
end = loop.time() + period
while loop.time() < end:
t = loop.time()
await asyncio.sleep(0.005)
lags.append(loop.time() - t - 0.005)
cur = cpu_stat()
LOOP_LAG_P99.set(sorted(lags)[int(len(lags) * 0.99)])
THROTTLED_PERIODS.inc(cur.get("nr_throttled", 0) - prev.get("nr_throttled", 0))
THROTTLED_SECONDS.inc((cur.get("throttled_usec", 0) - prev.get("throttled_usec", 0)) / 1e6)
prev = cur
The measurements separate two sources of lag that look alike from outside. With a busy thread and no limit, p99 lag was 10.6 ms and throttling zero: that is GIL contention, fixed by moving CPU work to processes. Under the 0.5-CPU quota, p99 jumped to 59.6 ms with throttling present: that is the quota, fixed by raising it or reducing CPU use. Lag without either points at blocking calls on the loop itself, as in finding blocking calls with asyncio debug mode.
Verify: during a load test that triggers throttling, loop-lag spikes and throttled-period increments occur in the same sampling windows.
3. Remove CPU work from latency-critical processes¶
The cause in the measurement was CPU work running in the same process as the event loop — in another thread, but sharing the process's quota and its GIL. Moving it out removes both effects:
from concurrent.futures import ProcessPoolExecutor
CPU_POOL = ProcessPoolExecutor(max_workers=1) # inside the same container: shares the quota
async def handle_report(request):
rows = await fetch_rows(request)
pdf = await asyncio.get_running_loop().run_in_executor(CPU_POOL, render_pdf, rows)
return Response(pdf, media_type="application/pdf")
A process pool in the same container removes GIL contention but still shares the container's quota, so heavy CPU work still throttles the event loop's process. For work that is both heavy and frequent, a separate deployment — a worker service consuming a queue, with its own CPU limit — isolates it completely; the job-queue patterns in Background Jobs & Task Queues apply. Keeping the latency-critical service's CPU use well below its limit is the goal; how far below is something the throttling metric tells you directly.
Verify: after moving CPU work out, the service's throttled-period rate falls to near zero at normal load.
4. Set limits that match the workload's bursts¶
Throttling happens when usage within a period exceeds the quota, even if the average is far lower. Async services are bursty — a request arrives, the loop runs flat out for a few milliseconds, then idles — so their limits need headroom above average usage:
resources:
requests:
cpu: "500m" # what the scheduler reserves: size from average use
limits:
cpu: "2" # the ceiling: size from bursts, or omit to avoid throttling
Requests drive scheduling and fair sharing under contention; limits drive throttling. A common production choice for latency-sensitive services is a request sized from typical usage and either a generous limit or no limit at all, relying on requests for fairness. Whatever the choice, the throttling metric from step 1 is the evidence: a service throttled during normal traffic needs a higher limit or less CPU work, whatever its average utilization says.
Verify: at peak normal traffic, throttled periods stay below a few percent of periods, and loop-lag p99 stays within budget.
5. Alert on throttling, not just on CPU utilization¶
CPU utilization averaged over a minute hides throttling completely — a service using 40% of its limit on average can be throttled in a third of its periods. Alert on the direct signal:
async def throttle_alert_check() -> None:
prev = cpu_stat()
while True:
await asyncio.sleep(60)
cur = cpu_stat()
periods = cur["nr_periods"] - prev["nr_periods"]
throttled = cur["nr_throttled"] - prev["nr_throttled"]
if periods and throttled / periods > 0.05:
log.warning("throttled in %.0f%% of CPU periods over the last minute",
100 * throttled / periods)
prev = cur
Container platforms usually export the same counters (container_cpu_cfs_throttled_periods_total and container_cpu_cfs_periods_total in cAdvisor-based monitoring), so the alert can live in the monitoring system instead; the in-process version has the advantage of sitting next to the loop-lag metric it explains. Pair it with the load-shedding approach in load shedding when the event loop is overloaded when throttling cannot be avoided during spikes.
Verify: a load test that pushes CPU above the limit fires the throttling alert, and normal traffic does not.
Verification¶
CPU throttling is visible and under control when:
cpu.statthrottling counters are exported next to event-loop lag.- Lag spikes are attributed to throttling, GIL contention or blocking calls by correlating the metrics.
- CPU-heavy work runs outside the latency-critical process, ideally outside its container.
- Limits leave room for bursts, and alerts fire on throttled periods, not average CPU.
Diagnostic Hook: plot loop-lag p99, throttled periods per minute and CPU usage on one chart per pod. Lag that tracks throttling at moderate average CPU is the quota; lag that tracks CPU usage without throttling is contention inside the process; lag with neither is a blocking call.
Pitfalls & edge cases¶
- Average CPU as the health signal. It hides throttling within 100 ms periods.
- Threads for CPU work in a quota-limited container. Measured: p99 lag 59.6 ms under 0.5 CPU.
- Process pools in the same container. They avoid the GIL but share the quota.
- Tight limits on bursty services. Size limits from bursts, requests from averages.
Frequently Asked Questions¶
What is CPU throttling in Kubernetes and Docker?
A CPU limit is a budget of CPU time per scheduling period (100 ms by default). When a container uses its budget early, all its threads are stopped until the next period. Under a 0.5-CPU limit, event-loop lag p99 rose to 59.6 ms in testing.
How do I detect CPU throttling from inside a container?
Read /sys/fs/cgroup/cpu.stat and track the change in nr_throttled and throttled_usec. In testing, 5 seconds under a 0.5-CPU limit recorded 51 throttled periods and 2.6 s of throttling.
Why is my asyncio service slow when CPU usage is low?
Bursts can exhaust the quota within a 100 ms period even when average usage is low, freezing the event loop until the next period. Compare loop lag with throttled periods.
Should latency-sensitive Python services have CPU limits?
Size requests from average usage and set limits with room for bursts, or omit them where the platform allows; then confirm with the throttling counters at peak load.
Related¶
- Containers & Serverless — up to the topic overview.
- Respecting container memory limits in async services — the other cgroup limit.
- Resilience, Cancellation & Error Handling — the section overview.