Skip to content

Load Testing & Benchmarking Async Services

Every capacity plan, timeout and alert threshold in an async service rests on measurements of latency and throughput, and those measurements are easy to get wrong in ways that flatter the system. Measured against a local aiohttp service doing 1 ms of CPU and 10 ms of I/O per request: a closed-loop load test with ten workers reported a p99 of 13 ms across a run that included a 1-second stall of the whole service, because its workers simply stopped sending while they waited; an open-loop generator sending 500 requests per second through the same stall reported p99 1,472 ms — what users would have seen. Averaging the p99s of four instances gave 299 ms when the true p99 of all requests was 81 ms. A microbenchmark that called asyncio.run for each measurement reported 44.0 µs for an operation that costs 3.15 µs. And the service held 13–14 ms latency up to about 940 requests per second, then crossed a knee: at 1,200 offered it served 933 with a median of 810 ms.

This section covers how to generate load, record latency, combine percentiles, benchmark small operations and find where a service saturates, without the errors above. The parent section, Resilience, Cancellation & Error Handling, uses these measurements to set timeouts, breakers and alerts.

Scope of this section:

  • Open-loop load generation that keeps sending when the service slows down.
  • Coordinated omission: what it hides, and how to measure around it.
  • Percentile arithmetic: merging histograms instead of averaging percentiles.
  • Microbenchmarks of coroutines and asyncio primitives with pyperf.
  • Finding the saturation point and the knee in the latency curve.

Architectural principles

  • Load arrives on its own schedule. Real users do not wait for your previous response before sending theirs; a load test must not either.
  • Latency is measured from the intended start. The clock starts when a request should have been sent, so time spent queued behind a stall counts.
  • Percentiles are computed, never averaged. Combine raw data or histograms, then take the percentile.
  • Benchmarks isolate what they measure. Event loop creation, process start-up and noisy neighbours stay outside the timed region or are averaged out by many runs.
  • Capacity is a curve, not a number. The useful result is latency versus offered load, with the knee marked.
A trustworthy load test, end to end A flow of 5 stages. A trustworthy load test, end to end open-loop schedule target rate measure from intended time no omission histograms per generator merge, then percentile never average repeat at rising rates find the knee Each step removes a way a test can report better numbers than users see.

Execution model: closed loops hide stalls

A closed-loop generator has a fixed number of workers, each sending a request, waiting for the response, and sending the next. When the service stalls, every worker is stuck waiting, so the generator stops generating; the stall shows up as a handful of slow samples among thousands of fast ones. Measured with ten workers and a 1-second stall in a 6-second run: the generator sent 4,287 requests instead of 5,145, the p99 stayed at 13 ms, and only the p99.9 (1,011 ms) hinted at the problem. An open-loop generator sends on a schedule regardless of responses:

async def open_loop(session, url: str, rate: float, seconds: float) -> list[float]:
    latencies: list[float] = []
    start = time.perf_counter()
    tasks = []

    async def one(intended: float) -> None:
        async with session.get(url) as response:
            await response.read()
        latencies.append(time.perf_counter() - intended)      # from the intended send time

    for i in range(int(rate * seconds)):
        intended = start + i / rate
        delay = intended - time.perf_counter()
        if delay > 0:
            await asyncio.sleep(delay)
        tasks.append(asyncio.create_task(one(intended)))
    await asyncio.gather(*tasks)
    return latencies

At 500 requests per second through the same 1-second stall, about 500 requests were affected and the p99 was 1,472 ms. The details — scheduling accuracy, connection limits, generator saturation — are in building an open-loop load generator in asyncio.

p99 reported for the same 1 s stall, by generator 3 horizontal bars comparing closed loop, 10 workers with the others. p99 reported for the same 1 s stall, by generator closed loop, 10 workers 13 ms closed loop, from intended time 987 ms open loop, 500 req/s 1,472 ms aiohttp service: 1 ms CPU + 10 ms I/O per request; 6 s runs; the stall froze the service for 1 s. The stall was identical; only the measurement differed.

Pattern catalogue

Measure from the intended start to avoid coordinated omission

When a generator's request is delayed because the previous one was slow, the delay is part of the user's experience. Recording latency from when the request should have been sent restores it. Measured: a closed loop with a fixed schedule per worker reported p99 13 ms per response but 987 ms from intended send times during the stall. See avoiding coordinated omission in latency benchmarks.

Merge histograms, then take percentiles

Percentiles do not average. Four instances, one with a slow dependency, had per-instance p99s of 76, 76, 77 and 969 ms; their average, 299 ms, matched nothing — the true p99 over all requests was 81 ms, and a p99 computed from merged histogram buckets was 86 ms. Averaging one-minute p99s over an hour gave 108 ms against a true 98 ms. See measuring latency percentiles without averaging them.

# PromQL: merge buckets across instances first, then the quantile
histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket[5m])))

Benchmark coroutines inside one long-lived loop

Microbenchmarks of async code must time the operation, not the event loop's creation. pyperf's bench_async_func runs many iterations inside one loop, in several worker processes, and reports mean and deviation. Measured: create_task plus await cost 3.15 µs ± 0.06, asyncio.sleep(0) 1.48 µs, awaiting a coroutine directly 48.6 ns; timeit with asyncio.run per call reported 44.0 µs for the same task operation. See microbenchmarking coroutines with pyperf and timeit.

Ramp the rate to find the knee

Run open-loop steps at increasing rates and record achieved throughput and latency percentiles at each. Measured: 13–14 ms latency up to 940 req/s; at 960 the median doubled to 32 ms; at 1,200 offered the service completed 933 per second with a median of 810 ms and a p99 of 1.73 s — past saturation, latency is set by the growing queue. See finding the saturation point of an async service.

Latency by offered load, one-process aiohttp service A grid of 5 rows by 4 columns. Latency by offered load, one-process aiohttp service offered req/s achieved p50 p99 600 598 12.2 ms 13.3 ms 900 896 13.2 ms 14.4 ms 940 935 14.1 ms 19.3 ms 960 952 31.8 ms 44.7 ms 1,200 933 809.5 ms 1,733.8 ms Latency measured from intended send time; 3 s per step.

Choosing what to run

Which measurement answers this question? A decision on What do you need to know with 4 outcomes. Which measurement answers this question? What do you need to know? latency at expected load open-loop test from intended time how much load it can take rate ramp knee of the curve cost of an operation pyperf microbenchmark one loop, many runs production percentiles merged histograms never averages Pick the measurement that matches the question, then avoid its known traps.

Load tests answer questions about the whole service under traffic; microbenchmarks answer questions about the cost of small pieces. Both are needed and neither substitutes for the other: a 3 µs operation can matter a great deal if it runs a thousand times per request, and a service can be slow for reasons no microbenchmark touches. Production metrics are the third leg — the same percentile rules apply when reading dashboards as when reading test reports. A useful habit is to check each load-test conclusion against production: if the test says the service saturates at 940 requests per second per process, production dashboards at peak should show loop busy fractions and latencies consistent with that, and a mismatch means the test's traffic mix or environment does not resemble reality.

Resource boundaries

  • Generator capacity: an asyncio generator is itself a single-threaded program; when its loop is saturated it sends late and its measurements absorb its own delays. Check that achieved rate matches target rate and that its own loop lag stays low; split across processes when it does not.
  • Connection limits: client pools cap concurrency (aiohttp defaults to 100 connections per session); a pool limit in the generator acts like a closed loop. Use limit=0 or a limit above the expected in-flight count.
  • Ports and descriptors: high rates without keep-alive exhaust ephemeral ports; reuse connections and raise ulimit -n.
  • Same-host effects: generator and service on one machine compete for CPU; fine for relative comparisons, misleading for absolute capacity.
  • Duration: short steps miss slow effects — GC, cache eviction, connection churn; run the final capacity test for minutes per step.

Integrated production example

A small harness that ramps an open-loop rate, records latency from intended times into histograms, and reports the knee:

import asyncio
import bisect
import time

import aiohttp

BOUNDS = [0.002 * 1.25 ** i for i in range(40)]         # ~2 ms to ~15 s, log-spaced


class Histogram:
    def __init__(self) -> None:
        self.counts = [0] * (len(BOUNDS) + 1)
        self.total = 0

    def record(self, seconds: float) -> None:
        self.counts[bisect.bisect_left(BOUNDS, seconds)] += 1
        self.total += 1

    def merge(self, other: "Histogram") -> None:
        self.counts = [a + b for a, b in zip(self.counts, other.counts)]
        self.total += other.total

    def quantile(self, q: float) -> float:
        target, cumulative = q * self.total, 0
        for i, count in enumerate(self.counts):
            cumulative += count
            if cumulative >= target:
                return BOUNDS[min(i, len(BOUNDS) - 1)]    # upper bound of the bucket
        return BOUNDS[-1]


async def step(session, url: str, rate: float, seconds: float) -> tuple[Histogram, float]:
    hist, tasks = Histogram(), []
    start = time.perf_counter()

    async def one(intended: float) -> None:
        try:
            async with session.get(url, timeout=aiohttp.ClientTimeout(total=30)) as r:
                await r.read()
        finally:
            hist.record(time.perf_counter() - intended)

    for i in range(int(rate * seconds)):
        intended = start + i / rate
        if (delay := intended - time.perf_counter()) > 0:
            await asyncio.sleep(delay)
        tasks.append(asyncio.create_task(one(intended)))
    await asyncio.gather(*tasks, return_exceptions=True)
    return hist, hist.total / (time.perf_counter() - start)


async def ramp(url: str, rates: list[float], seconds: float = 60.0, slo_p99: float = 0.05) -> None:
    connector = aiohttp.TCPConnector(limit=0)
    async with aiohttp.ClientSession(connector=connector) as session:
        for rate in rates:
            hist, achieved = await step(session, url, rate, seconds)
            p50, p99 = hist.quantile(0.5), hist.quantile(0.99)
            flag = "  <- over SLO" if p99 > slo_p99 else ""
            print(f"offered {rate:7.0f}/s  achieved {achieved:7.0f}/s  p50 {p50*1e3:7.1f} ms  "
                  f"p99 {p99*1e3:7.1f} ms{flag}")
            await asyncio.sleep(5)                               # let queues drain between steps

Log-spaced buckets keep relative precision constant from milliseconds to seconds and make histograms from several generator processes mergeable by adding counts. Reporting achieved against offered rate exposes both saturation (achieved stops rising) and a generator that cannot keep up. The SLO flag turns the curve into a capacity number: the highest rate whose p99 stays within the objective.

Shaping realistic load

A fixed interval between requests is the simplest open-loop schedule and the least realistic. Real arrivals are bursty: independent users produce a Poisson process, in which gaps are exponentially distributed and short bursts are normal. Bursts matter for async services because they fill queues and connection pools momentarily even when the average rate is comfortable:

import random

def poisson_schedule(rate: float, seconds: float, seed: int = 1):
    rng = random.Random(seed)
    t = 0.0
    while t < seconds:
        t += rng.expovariate(rate)          # exponential gaps: bursts and lulls
        yield t

Seeding the generator makes runs repeatable while keeping the burstiness. Beyond timing, realistic load needs a realistic mix: the proportion of each endpoint, payload sizes drawn from production, cache hit rates similar to production (a test that hits the same key every time measures the cache, not the service), and authentication paths exercised. Warm the service up before measuring — the first seconds include connection-pool growth, JIT-like caches in libraries and cold OS page caches — and discard that period from the results.

Reporting results people can trust

A number without context invites misreading. Every load-test or benchmark report should state what was measured and how, so that a later run can be compared with it:

  • Versions and environment: Python version, library versions, hardware or instance type, whether generator and service shared a machine.
  • Method: open or closed loop, schedule (fixed or Poisson), measured from intended or actual send time, duration and warm-up per step.
  • Results as distributions: p50, p90, p99 and p99.9 with counts, plus achieved versus offered rate and error counts — not a single average.
  • Variance: at least three repetitions per configuration, with the spread; a 5% difference inside a 10% run-to-run spread is noise.

Comparisons are most reliable when only one thing changes between runs, on the same machine, close together in time. For code changes, run old and new versions alternately rather than all old runs followed by all new ones, so slow drifts in the environment affect both equally. Keep raw histograms, not just summary numbers; they allow re-analysis when a new question comes up.

Load testing in CI

Full capacity tests take minutes per step and need stable hardware, so they rarely belong in every CI run. A smaller, regular check still catches regressions early: run a short open-loop test at a fixed, moderate rate against each build on the same runner type, and compare its p99 and CPU per request with the previous build's. Fail or flag the build when either moves by more than the measured run-to-run spread. Microbenchmarks of hot paths fit the same pattern — pyperf can compare two result files and report whether a difference is significant — and they run in seconds. Reserve the full ramp for releases and infrastructure changes, where the absolute capacity number is what matters.

Diagnostic hook callout

Diagnostic Hook: with every load-test report, record the generator's achieved-versus-target rate, its own event-loop lag, and whether latency was measured from intended or actual send times. A report without those three facts cannot be trusted: a lagging generator, a closed loop, or response-time-only latency each make the service look better than it is.

Failure modes

  • Closed-loop generators. Measured: p99 of 13 ms across a 1-second stall.
  • Averaging percentiles. Measured: 299 ms reported for a true 81 ms.
  • Timing event loop creation. Measured: 44.0 µs reported for a 3.15 µs operation.
  • A single "max throughput" number. The knee and the latency at each rate are what matter.
  • An overloaded generator. It sends late and hides the service's latency in its own.

Frequently Asked Questions

What is the difference between open-loop and closed-loop load testing?

A closed loop has fixed workers that wait for each response before sending again, so it slows down when the service does. An open loop sends on a schedule regardless. In testing, a closed loop reported p99 13 ms through a 1-second stall that an open loop reported as 1,472 ms.

What is coordinated omission?

The measurement error where a load generator stops sending while waiting for slow responses, so the delays users would have experienced are never recorded. Measuring latency from intended send times corrects it.

Can I average p99 latency across servers?

No. Merge the raw data or histogram buckets first and compute the percentile from the result; averaging four instances' p99s gave 299 ms when the true p99 was 81 ms in testing.

How do I benchmark asyncio code accurately?

Run many iterations inside one event loop, in several processes, with a tool such as pyperf's bench_async_func. Creating a loop per measurement made create_task appear 14 times more expensive in testing.

How do I find the maximum load an async service can handle?

Ramp an open-loop rate in steps and plot achieved throughput and p99 at each; the capacity is the highest rate whose p99 meets your objective, below the knee where latency starts to climb.