Building an Open-Loop Load Generator in asyncio¶
A load generator decides what a load test can see. The common design — a fixed number of workers, each sending a request and waiting for the answer before sending the next — slows down whenever the service does, so slowdowns produce fewer requests and fewer slow samples. An open-loop generator sends on a schedule that does not depend on responses, the way independent users do. Measured against a local aiohttp service (1 ms CPU and 10 ms I/O per request) with a 1-second stall in a 6-second run: ten closed-loop workers sent 4,287 requests — 858 fewer than without the stall — and reported a p99 of 13 ms. An open-loop generator at 500 requests per second sent all 3,000, about 500 of them during or behind the stall, and reported a p99 of 1,472 ms. Without the stall, both reported 12–13 ms. This guide builds the open-loop generator, keeps it from becoming the bottleneck, and makes its output trustworthy.
Prerequisites¶
- Python 3.11+,
pip install aiohttp(the client used for the measurements). - Why this matters, from Load Testing & Benchmarking.
- Client connection limits, from configuring aiohttp TCPConnector limits.
1. Schedule sends by time, not by responses¶
Compute each request's intended send time from the start time and the rate, sleep until then, and launch the request as its own task so the loop never waits for responses:
import asyncio
import time
import aiohttp
async def open_loop(url: str, rate: float, seconds: float) -> list[float]:
latencies: list[float] = []
connector = aiohttp.TCPConnector(limit=0) # no client-side concurrency cap
async with aiohttp.ClientSession(connector=connector) as session:
start = time.perf_counter()
tasks = []
async def one(intended: float) -> None:
async with session.get(url) as response:
await response.read()
latencies.append(time.perf_counter() - intended)
for i in range(int(rate * seconds)):
intended = start + i / rate
delay = intended - time.perf_counter()
if delay > 0:
await asyncio.sleep(delay)
tasks.append(asyncio.create_task(one(intended)))
await asyncio.gather(*tasks, return_exceptions=True)
return latencies
Measured: 3,000 requests at 500 per second regardless of a 1-second service stall, and a p99 of 1,472 ms that reflected the stall. Deriving each send time from start + i / rate rather than sleeping a fixed interval after each send keeps the schedule from drifting when the loop is briefly busy. limit=0 on the connector matters: a pool limit would make requests queue inside the client, turning the generator back into a closed loop.
Verify: the number of requests sent equals rate × duration, whatever the service does.
2. Record latency from the intended send time¶
Latency should start when the request was supposed to be sent. If the generator itself falls behind — its loop busy, the machine loaded — requests go out late, and measuring from the actual send hides that delay:
async def one(intended: float) -> None:
sent = time.perf_counter()
async with session.get(url) as response:
await response.read()
done = time.perf_counter()
service_latency.append(done - sent) # what the server took
user_latency.append(done - intended) # what a user arriving on schedule saw
send_lag.append(sent - intended) # how late the generator was
Keep all three. user_latency is the number to report; send_lag tells you whether the generator kept up — it should stay near zero; if it grows, the generator is the bottleneck and its results are suspect. In the measured open-loop runs the two latency numbers agreed (p99 1,472 vs 1,473 ms), confirming the generator was on schedule. The topic of why this matters for slow responses is covered in avoiding coordinated omission in latency benchmarks.
Verify: the p99 of send_lag is a few milliseconds at most for every run you report.
3. Keep the generator from becoming the bottleneck¶
An asyncio generator is one thread. Each request costs it parsing, task creation and bookkeeping; at high rates it saturates before the service does:
async def watch_generator(stop: asyncio.Event, lag_samples: list[float]) -> None:
loop = asyncio.get_running_loop()
while not stop.is_set():
t = loop.time()
await asyncio.sleep(0.01)
lag_samples.append(loop.time() - t - 0.01) # the generator's own loop lag
# For rates beyond one process, run N generator processes, each at rate / N, and merge histograms
Monitor the generator's own loop lag alongside send lag; if either grows, split the load across processes (multiprocessing, or several machines) and merge their histograms afterwards, as in measuring latency percentiles without averaging them. The choice of HTTP client matters: in testing elsewhere on this site, aiohttp sustained far higher request rates per process than httpx at high concurrency, as measured in choosing between httpx and aiohttp.
Verify: at your highest test rate, the generator's loop lag p99 is below a few milliseconds.
4. Use realistic arrival patterns¶
Uniform spacing is easy to reason about but smoother than real traffic. Independent users produce Poisson arrivals — exponentially distributed gaps — which include short bursts:
import random
def poisson_times(rate: float, seconds: float, seed: int = 1):
rng = random.Random(seed)
t = 0.0
while True:
t += rng.expovariate(rate)
if t >= seconds:
return
yield t
async def run_schedule(session, url, offsets, record):
start = time.perf_counter()
tasks = []
for offset in offsets:
intended = start + offset
if (delay := intended - time.perf_counter()) > 0:
await asyncio.sleep(delay)
tasks.append(asyncio.create_task(send(session, url, intended, record)))
await asyncio.gather(*tasks, return_exceptions=True)
Bursts fill queues and pools briefly even at a comfortable average rate, so Poisson arrivals give higher tail latency than uniform ones at the same mean — closer to production. A seeded generator keeps runs repeatable. For more realism, replay a captured production trace: its timestamps become the offsets, and its request mix becomes the URLs.
Verify: a Poisson run at the same mean rate shows equal or higher p99 than the uniform run.
5. Bound the run and report it fully¶
An open-loop generator has no natural limit on in-flight requests: if the service stops answering, tasks accumulate. Bound each request with a timeout, and report every dimension of the run:
async def send(session, url, intended, record) -> None:
try:
async with session.get(url, timeout=aiohttp.ClientTimeout(total=30)) as response:
await response.read()
record(intended, time.perf_counter(), response.status)
except (aiohttp.ClientError, TimeoutError) as exc:
record(intended, time.perf_counter(), type(exc).__name__)
def report(rate, duration, results) -> dict:
ok = [done - intended for intended, done, status in results if status == 200]
return {
"offered_rps": rate, "achieved_rps": len(ok) / duration,
"p50_ms": pct(ok, 0.50) * 1e3, "p99_ms": pct(ok, 0.99) * 1e3, "p999_ms": pct(ok, 0.999) * 1e3,
"errors": sum(1 for *_, s in results if s != 200),
}
The per-request timeout caps memory and file descriptors when the service hangs; count timeouts as errors rather than dropping them, or the report omits exactly the worst outcomes. Report offered and achieved rates together: when they diverge, the service is saturated (or erroring), which is the subject of finding the saturation point of an async service.
Verify: a run against a service that stops responding ends at the timeout, reports the failures, and does not exhaust the generator's memory.
Verification¶
An open-loop generator is trustworthy when:
- Send times come from the schedule, and the client pool does not cap concurrency.
- Latency is measured from intended send times, with send lag reported.
- The generator's own loop lag stays low, or load is split across processes.
- Every request has a timeout, and errors are part of the report.
Diagnostic Hook: plot send lag against time for each run. A flat line near zero means the generator kept its schedule; spikes that coincide with latency spikes mean the generator, not the service, may be responsible — rerun with the load split across more processes before drawing conclusions.
Pitfalls & edge cases¶
- Closed-loop workers. Measured: 858 fewer requests and a p99 of 13 ms through a 1 s stall.
- Client pool limits. They turn an open loop back into a closed one.
- Measuring from the actual send time. A late generator hides its own delay.
- No per-request timeout. A hung service exhausts the generator.
Frequently Asked Questions¶
How do I write an open-loop load generator in Python asyncio?
Compute each request's intended send time from the start time and rate, sleep until it, launch the request with asyncio.create_task without awaiting it, and record latency from the intended time. In testing it kept sending 500 req/s through a 1 s stall.
Why does my load test show low latency when the service had a pause?
A closed-loop generator stops sending while its workers wait, so few requests experience the pause. Ten workers reported p99 13 ms across a 1 s stall that an open loop reported as 1,472 ms.
How do I know my asyncio load generator is keeping up?
Track send lag, the difference between actual and intended send times, and the generator's own event-loop lag; both should stay near zero.
Should load tests use Poisson arrivals?
For simulating independent users, yes: exponentially distributed gaps include realistic bursts that uniform spacing lacks.
Related¶
- Load Testing & Benchmarking — up to the topic overview.
- Avoiding coordinated omission in latency benchmarks — the measurement error open loops avoid.
- Resilience, Cancellation & Error Handling — the section overview.