Load Testing & Benchmarking Async Services¶
Every capacity plan, timeout and alert threshold in an async service rests on measurements of latency and throughput, and those measurements are easy to get wrong in ways that flatter the system. Measured against a local aiohttp service doing 1 ms of CPU and 10 ms of I/O per request: a closed-loop load test with ten workers reported a p99 of 13 ms across a run that included a 1-second stall of the whole service, because its workers simply stopped sending while they waited; an open-loop generator sending 500 requests per second through the same stall reported p99 1,472 ms — what users would have seen. Averaging the p99s of four instances gave 299 ms when the true p99 of all requests was 81 ms. A microbenchmark that called asyncio.run for each measurement reported 44.0 µs for an operation that costs 3.15 µs. And the service held 13–14 ms latency up to about 940 requests per second, then crossed a knee: at 1,200 offered it served 933 with a median of 810 ms.
This section covers how to generate load, record latency, combine percentiles, benchmark small operations and find where a service saturates, without the errors above. The parent section, Resilience, Cancellation & Error Handling, uses these measurements to set timeouts, breakers and alerts.
Scope of this section:
- Open-loop load generation that keeps sending when the service slows down.
- Coordinated omission: what it hides, and how to measure around it.
- Percentile arithmetic: merging histograms instead of averaging percentiles.
- Microbenchmarks of coroutines and asyncio primitives with pyperf.
- Finding the saturation point and the knee in the latency curve.
Architectural principles¶
- Load arrives on its own schedule. Real users do not wait for your previous response before sending theirs; a load test must not either.
- Latency is measured from the intended start. The clock starts when a request should have been sent, so time spent queued behind a stall counts.
- Percentiles are computed, never averaged. Combine raw data or histograms, then take the percentile.
- Benchmarks isolate what they measure. Event loop creation, process start-up and noisy neighbours stay outside the timed region or are averaged out by many runs.
- Capacity is a curve, not a number. The useful result is latency versus offered load, with the knee marked.
Execution model: closed loops hide stalls¶
A closed-loop generator has a fixed number of workers, each sending a request, waiting for the response, and sending the next. When the service stalls, every worker is stuck waiting, so the generator stops generating; the stall shows up as a handful of slow samples among thousands of fast ones. Measured with ten workers and a 1-second stall in a 6-second run: the generator sent 4,287 requests instead of 5,145, the p99 stayed at 13 ms, and only the p99.9 (1,011 ms) hinted at the problem. An open-loop generator sends on a schedule regardless of responses:
async def open_loop(session, url: str, rate: float, seconds: float) -> list[float]:
latencies: list[float] = []
start = time.perf_counter()
tasks = []
async def one(intended: float) -> None:
async with session.get(url) as response:
await response.read()
latencies.append(time.perf_counter() - intended) # from the intended send time
for i in range(int(rate * seconds)):
intended = start + i / rate
delay = intended - time.perf_counter()
if delay > 0:
await asyncio.sleep(delay)
tasks.append(asyncio.create_task(one(intended)))
await asyncio.gather(*tasks)
return latencies
At 500 requests per second through the same 1-second stall, about 500 requests were affected and the p99 was 1,472 ms. The details — scheduling accuracy, connection limits, generator saturation — are in building an open-loop load generator in asyncio.
Pattern catalogue¶
Measure from the intended start to avoid coordinated omission¶
When a generator's request is delayed because the previous one was slow, the delay is part of the user's experience. Recording latency from when the request should have been sent restores it. Measured: a closed loop with a fixed schedule per worker reported p99 13 ms per response but 987 ms from intended send times during the stall. See avoiding coordinated omission in latency benchmarks.
Merge histograms, then take percentiles¶
Percentiles do not average. Four instances, one with a slow dependency, had per-instance p99s of 76, 76, 77 and 969 ms; their average, 299 ms, matched nothing — the true p99 over all requests was 81 ms, and a p99 computed from merged histogram buckets was 86 ms. Averaging one-minute p99s over an hour gave 108 ms against a true 98 ms. See measuring latency percentiles without averaging them.
# PromQL: merge buckets across instances first, then the quantile
histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket[5m])))
Benchmark coroutines inside one long-lived loop¶
Microbenchmarks of async code must time the operation, not the event loop's creation. pyperf's bench_async_func runs many iterations inside one loop, in several worker processes, and reports mean and deviation. Measured: create_task plus await cost 3.15 µs ± 0.06, asyncio.sleep(0) 1.48 µs, awaiting a coroutine directly 48.6 ns; timeit with asyncio.run per call reported 44.0 µs for the same task operation. See microbenchmarking coroutines with pyperf and timeit.
Ramp the rate to find the knee¶
Run open-loop steps at increasing rates and record achieved throughput and latency percentiles at each. Measured: 13–14 ms latency up to 940 req/s; at 960 the median doubled to 32 ms; at 1,200 offered the service completed 933 per second with a median of 810 ms and a p99 of 1.73 s — past saturation, latency is set by the growing queue. See finding the saturation point of an async service.
Choosing what to run¶
Load tests answer questions about the whole service under traffic; microbenchmarks answer questions about the cost of small pieces. Both are needed and neither substitutes for the other: a 3 µs operation can matter a great deal if it runs a thousand times per request, and a service can be slow for reasons no microbenchmark touches. Production metrics are the third leg — the same percentile rules apply when reading dashboards as when reading test reports. A useful habit is to check each load-test conclusion against production: if the test says the service saturates at 940 requests per second per process, production dashboards at peak should show loop busy fractions and latencies consistent with that, and a mismatch means the test's traffic mix or environment does not resemble reality.
Resource boundaries¶
- Generator capacity: an asyncio generator is itself a single-threaded program; when its loop is saturated it sends late and its measurements absorb its own delays. Check that achieved rate matches target rate and that its own loop lag stays low; split across processes when it does not.
- Connection limits: client pools cap concurrency (aiohttp defaults to 100 connections per session); a pool limit in the generator acts like a closed loop. Use
limit=0or a limit above the expected in-flight count. - Ports and descriptors: high rates without keep-alive exhaust ephemeral ports; reuse connections and raise
ulimit -n. - Same-host effects: generator and service on one machine compete for CPU; fine for relative comparisons, misleading for absolute capacity.
- Duration: short steps miss slow effects — GC, cache eviction, connection churn; run the final capacity test for minutes per step.
Integrated production example¶
A small harness that ramps an open-loop rate, records latency from intended times into histograms, and reports the knee:
import asyncio
import bisect
import time
import aiohttp
BOUNDS = [0.002 * 1.25 ** i for i in range(40)] # ~2 ms to ~15 s, log-spaced
class Histogram:
def __init__(self) -> None:
self.counts = [0] * (len(BOUNDS) + 1)
self.total = 0
def record(self, seconds: float) -> None:
self.counts[bisect.bisect_left(BOUNDS, seconds)] += 1
self.total += 1
def merge(self, other: "Histogram") -> None:
self.counts = [a + b for a, b in zip(self.counts, other.counts)]
self.total += other.total
def quantile(self, q: float) -> float:
target, cumulative = q * self.total, 0
for i, count in enumerate(self.counts):
cumulative += count
if cumulative >= target:
return BOUNDS[min(i, len(BOUNDS) - 1)] # upper bound of the bucket
return BOUNDS[-1]
async def step(session, url: str, rate: float, seconds: float) -> tuple[Histogram, float]:
hist, tasks = Histogram(), []
start = time.perf_counter()
async def one(intended: float) -> None:
try:
async with session.get(url, timeout=aiohttp.ClientTimeout(total=30)) as r:
await r.read()
finally:
hist.record(time.perf_counter() - intended)
for i in range(int(rate * seconds)):
intended = start + i / rate
if (delay := intended - time.perf_counter()) > 0:
await asyncio.sleep(delay)
tasks.append(asyncio.create_task(one(intended)))
await asyncio.gather(*tasks, return_exceptions=True)
return hist, hist.total / (time.perf_counter() - start)
async def ramp(url: str, rates: list[float], seconds: float = 60.0, slo_p99: float = 0.05) -> None:
connector = aiohttp.TCPConnector(limit=0)
async with aiohttp.ClientSession(connector=connector) as session:
for rate in rates:
hist, achieved = await step(session, url, rate, seconds)
p50, p99 = hist.quantile(0.5), hist.quantile(0.99)
flag = " <- over SLO" if p99 > slo_p99 else ""
print(f"offered {rate:7.0f}/s achieved {achieved:7.0f}/s p50 {p50*1e3:7.1f} ms "
f"p99 {p99*1e3:7.1f} ms{flag}")
await asyncio.sleep(5) # let queues drain between steps
Log-spaced buckets keep relative precision constant from milliseconds to seconds and make histograms from several generator processes mergeable by adding counts. Reporting achieved against offered rate exposes both saturation (achieved stops rising) and a generator that cannot keep up. The SLO flag turns the curve into a capacity number: the highest rate whose p99 stays within the objective.
Shaping realistic load¶
A fixed interval between requests is the simplest open-loop schedule and the least realistic. Real arrivals are bursty: independent users produce a Poisson process, in which gaps are exponentially distributed and short bursts are normal. Bursts matter for async services because they fill queues and connection pools momentarily even when the average rate is comfortable:
import random
def poisson_schedule(rate: float, seconds: float, seed: int = 1):
rng = random.Random(seed)
t = 0.0
while t < seconds:
t += rng.expovariate(rate) # exponential gaps: bursts and lulls
yield t
Seeding the generator makes runs repeatable while keeping the burstiness. Beyond timing, realistic load needs a realistic mix: the proportion of each endpoint, payload sizes drawn from production, cache hit rates similar to production (a test that hits the same key every time measures the cache, not the service), and authentication paths exercised. Warm the service up before measuring — the first seconds include connection-pool growth, JIT-like caches in libraries and cold OS page caches — and discard that period from the results.
Reporting results people can trust¶
A number without context invites misreading. Every load-test or benchmark report should state what was measured and how, so that a later run can be compared with it:
- Versions and environment: Python version, library versions, hardware or instance type, whether generator and service shared a machine.
- Method: open or closed loop, schedule (fixed or Poisson), measured from intended or actual send time, duration and warm-up per step.
- Results as distributions: p50, p90, p99 and p99.9 with counts, plus achieved versus offered rate and error counts — not a single average.
- Variance: at least three repetitions per configuration, with the spread; a 5% difference inside a 10% run-to-run spread is noise.
Comparisons are most reliable when only one thing changes between runs, on the same machine, close together in time. For code changes, run old and new versions alternately rather than all old runs followed by all new ones, so slow drifts in the environment affect both equally. Keep raw histograms, not just summary numbers; they allow re-analysis when a new question comes up.
Load testing in CI¶
Full capacity tests take minutes per step and need stable hardware, so they rarely belong in every CI run. A smaller, regular check still catches regressions early: run a short open-loop test at a fixed, moderate rate against each build on the same runner type, and compare its p99 and CPU per request with the previous build's. Fail or flag the build when either moves by more than the measured run-to-run spread. Microbenchmarks of hot paths fit the same pattern — pyperf can compare two result files and report whether a difference is significant — and they run in seconds. Reserve the full ramp for releases and infrastructure changes, where the absolute capacity number is what matters.
Diagnostic hook callout¶
Diagnostic Hook: with every load-test report, record the generator's achieved-versus-target rate, its own event-loop lag, and whether latency was measured from intended or actual send times. A report without those three facts cannot be trusted: a lagging generator, a closed loop, or response-time-only latency each make the service look better than it is.
Failure modes¶
- Closed-loop generators. Measured: p99 of 13 ms across a 1-second stall.
- Averaging percentiles. Measured: 299 ms reported for a true 81 ms.
- Timing event loop creation. Measured: 44.0 µs reported for a 3.15 µs operation.
- A single "max throughput" number. The knee and the latency at each rate are what matter.
- An overloaded generator. It sends late and hides the service's latency in its own.
Frequently Asked Questions¶
What is the difference between open-loop and closed-loop load testing?
A closed loop has fixed workers that wait for each response before sending again, so it slows down when the service does. An open loop sends on a schedule regardless. In testing, a closed loop reported p99 13 ms through a 1-second stall that an open loop reported as 1,472 ms.
What is coordinated omission?
The measurement error where a load generator stops sending while waiting for slow responses, so the delays users would have experienced are never recorded. Measuring latency from intended send times corrects it.
Can I average p99 latency across servers?
No. Merge the raw data or histogram buckets first and compute the percentile from the result; averaging four instances' p99s gave 299 ms when the true p99 was 81 ms in testing.
How do I benchmark asyncio code accurately?
Run many iterations inside one event loop, in several processes, with a tool such as pyperf's bench_async_func. Creating a loop per measurement made create_task appear 14 times more expensive in testing.
How do I find the maximum load an async service can handle?
Ramp an open-loop rate in steps and plot achieved throughput and p99 at each; the capacity is the highest rate whose p99 meets your objective, below the knee where latency starts to climb.
Related¶
- Building an open-loop load generator in asyncio — the generator, step by step.
- Avoiding coordinated omission in latency benchmarks — measuring what users experience.
- Measuring latency percentiles without averaging them — percentile arithmetic done right.
- Microbenchmarking coroutines with pyperf and timeit — the cost of small operations.
- Finding the saturation point of an async service — the knee of the curve.
- Alerting on event loop saturation — the production signal of the same knee.
- Resilience, Cancellation & Error Handling — up to the section overview.