Finding the Saturation Point of an Async Service¶
Every service has a rate beyond which it cannot keep up: requests arrive faster than they are completed, queues grow, and latency is no longer set by the work but by the wait. For a single-threaded event loop, that point is where the loop is busy all the time. Finding it — and the knee just before it where tail latency starts to rise — gives you the capacity number for planning, the threshold for load shedding and the basis for alerts. Ramped with an open-loop generator against a one-process aiohttp service doing 1 ms of CPU and 10 ms of I/O per request: latency stayed at 13–14 ms (p99 13–19 ms) from 200 up to 940 requests per second; at 960 the median doubled to 32 ms and the p99 reached 45 ms; at 1,000 offered the service completed only 934 per second with a 212 ms median; at 1,200 it completed 933 with a median of 810 ms and a p99 of 1.73 s. This guide runs the ramp, reads the curve, and turns it into capacity and protection settings.
Prerequisites¶
- Python 3.11+,
pip install aiohttpfor the generator. - An open-loop generator, from building an open-loop load generator in asyncio.
- Loop saturation metrics, from alerting on event loop saturation.
1. Ramp an open-loop rate in steps¶
Offer a fixed rate for a while, measure, pause, and step up. Measure latency from intended send times and record the achieved rate:
async def step(session, url: str, rate: float, seconds: float = 30.0) -> dict:
latencies, tasks, errors = [], [], 0
start = time.perf_counter()
async def one(intended: float) -> None:
nonlocal errors
try:
async with session.get(url, timeout=aiohttp.ClientTimeout(total=10)) as r:
await r.read()
latencies.append(time.perf_counter() - intended)
except (aiohttp.ClientError, TimeoutError):
errors += 1
for i in range(int(rate * seconds)):
intended = start + i / rate
if (delay := intended - time.perf_counter()) > 0:
await asyncio.sleep(delay)
tasks.append(asyncio.create_task(one(intended)))
await asyncio.gather(*tasks)
elapsed = time.perf_counter() - start
return {"offered": rate, "achieved": len(latencies) / elapsed, "errors": errors,
"p50": pct(latencies, 0.50), "p99": pct(latencies, 0.99)}
for rate in (200, 400, 600, 800, 900, 940, 960, 1000, 1200):
print(await step(session, URL, rate))
await asyncio.sleep(5) # let queues drain between steps
The measurements in this guide used 3-second steps for speed; for real capacity numbers, use at least 30–60 seconds per step so that slow effects — garbage collection, pool growth, cache churn — appear. Run the generator on a different machine, or at least confirm it is not saturated itself, and keep its client pool unbounded so it stays open-loop.
Verify: at low rates, achieved equals offered and latency equals the service's unloaded latency.
2. Read the curve: flat, knee, wall¶
The measured curve has the three regions every single-resource system shows:
offered achieved p50 p99
200 200 12.4 ms 13.0 ms flat: latency = service time
800 797 12.0 ms 13.2 ms
940 935 14.1 ms 19.3 ms knee: tail rises first
960 952 31.8 ms 44.7 ms
1000 934 211.8 ms 539.7 ms wall: achieved stops rising, queue grows
1200 933 809.5 ms 1733.8 ms
In the flat region, latency is the work itself (11 ms of service time plus overhead). Near the knee, requests start to wait for the loop: the p99 rises first, then the median. Past saturation — about 935–950 requests per second here, set by roughly 1.05 ms of loop time per request — achieved throughput stops rising and every extra request offered only lengthens the queue, so latency grows with how long the overload lasts. Note that the achieved rate at 1,200 offered (933) is slightly lower than at 960 offered (952): overload costs efficiency.
Verify: your curve shows where achieved rate stops tracking offered rate; that is the saturation throughput.
3. Turn the curve into a capacity number¶
Capacity is not the saturation throughput; it is the highest rate at which latency still meets your objective, with headroom:
def capacity(results: list[dict], slo_p99: float, headroom: float = 0.7) -> float:
"""Highest offered rate meeting the p99 objective, times a headroom factor."""
ok = [r["offered"] for r in results if r["p99"] <= slo_p99 and r["errors"] == 0]
return max(ok) * headroom if ok else 0.0
capacity(measured, slo_p99=0.025) # p99 <= 25 ms met up to 940 -> plan for ~660 req/s per process
With a p99 objective of 25 ms, the measured service met it up to 940 requests per second; planning at about 70% of that leaves room for bursts (real arrivals are not uniform), slower instances, and gradual code changes that make each request a little more expensive. Divide expected peak traffic by the per-process capacity to size the fleet, and re-measure after significant changes. The per-request loop time (here about 1 ms of CPU) is the lever: halving it roughly doubles capacity.
Verify: the documented capacity per process comes from a recorded ramp with its objective and headroom stated.
4. Find what saturates¶
Knowing the rate is half the result; knowing which resource sets it tells you how to raise it. Observe the service during the ramp:
# During each step, sample on the service:
# loop busy fraction / loop lag -> the event loop (CPU on one core)
# process CPU vs cores -> CPU overall
# pool wait times -> database or HTTP client pools
# queue depths -> internal queues and backpressure points
# dependency latency -> something downstream
For the measured service the event loop saturated: one core busy with about 1 ms of CPU per request. The fix for that is more processes (one loop per core) or less CPU per request. If instead pool waits rise at the knee while the loop is idle, the bottleneck is the pool or the dependency behind it; if dependency latency rises first, the limit is downstream and adding processes will not help. A profile at the knee, as in continuous profiling of async services, shows where the loop's time goes.
Verify: at the knee, exactly one resource's utilization is near its limit; if none is, the generator may be the bottleneck.
5. Protect the service beyond the knee¶
Past saturation, every extra request makes all requests slower — measured, the median went from 14 ms to 810 ms for a 28% increase in offered load. Shed load before that happens, using the numbers from the ramp:
class AdmissionControl:
"""Reject early when in-flight work exceeds what the measured capacity allows."""
def __init__(self, capacity_rps: float, target_latency_s: float) -> None:
self.limit = max(1, int(capacity_rps * target_latency_s * 1.5)) # Little's law + margin
self.in_flight = 0
def try_enter(self) -> bool:
if self.in_flight >= self.limit:
return False
self.in_flight += 1
return True
def leave(self) -> None:
self.in_flight -= 1
# capacity 940 req/s, target 15 ms -> about 21 requests in flight per process
Little's law gives the in-flight count at the knee — throughput × latency, about 940 × 0.015 ≈ 14 requests — and a cap a little above it keeps the service on the flat side of the curve under overload, rejecting the excess quickly with 503 instead of making everyone wait. The broader techniques are in load shedding when the event loop is overloaded. Then re-run the ramp: with admission control, latency for admitted requests should stay near the knee value even at 1,200 offered.
Verify: a ramp past saturation with admission control shows bounded latency for successful requests and a rising share of fast rejections.
Verification¶
The saturation point is understood when:
- An open-loop ramp records offered rate, achieved rate and percentiles at each step.
- The knee and the saturation throughput are identified, and capacity is set below the knee with headroom.
- The saturating resource is known, so scaling decisions target it.
- Admission control keeps the service on the flat side under overload, verified by a re-run.
Diagnostic Hook: keep the ramp results for each release and plot capacity per process over time. A capacity that drifts downward release by release is the cumulative cost of small per-request additions — the kind of regression no single microbenchmark catches.
Pitfalls & edge cases¶
- Reporting the wall as capacity. At 1,000 offered, achieved was 934 with a 212 ms median.
- Closed-loop ramps. They never overload the service and never find the knee.
- Short steps. Slow effects appear only after minutes.
- Scaling the wrong resource. More processes do not help if a dependency saturates first.
Frequently Asked Questions¶
How do I find the maximum throughput of an asyncio service?
Ramp an open-loop request rate in steps, recording achieved rate and latency percentiles at each. In testing, a one-process aiohttp service tracked the offered rate up to about 950 req/s and then stayed near 933 while latency grew.
What is the knee in a latency-versus-load curve?
The point where latency starts to rise because requests begin to wait for a saturated resource; the p99 rises first, then the median. In testing it lay between 940 and 960 req/s.
How much headroom should capacity planning leave?
Plan at a fraction of the highest rate that meets your latency objective, often around 70%, to absorb bursts, slower instances and gradual cost increases.
How do I keep a service fast when it is overloaded?
Cap in-flight requests near throughput times target latency (Little's law) and reject the excess quickly, so admitted requests stay on the flat part of the curve.
Related¶
- Load Testing & Benchmarking — up to the topic overview.
- Microbenchmarking coroutines with pyperf and timeit — reducing the per-request cost that sets the knee.
- Resilience, Cancellation & Error Handling — the section overview.