Avoiding Coordinated Omission in Latency Benchmarks¶
Coordinated omission is the measurement error where a load generator, waiting for a slow response, does not send the requests it would have sent — so the delays those requests would have suffered are never recorded. The generator "coordinates" with the system under test and omits exactly the bad samples. Measured against a local aiohttp service with a 1-second stall in a 6-second run: a closed-loop generator with ten workers reported p50 11 ms and p99 13 ms — the stall appeared only in the p99.9 (1,011 ms), carried by ten samples. The same closed loop, given a fixed schedule per worker and measuring each request from its intended send time, reported p99 987 ms for an identical stall. An open-loop generator at 500 requests per second reported p99 1,472 ms. Without the stall, all three reported 13 ms or less. This guide shows how the error arises, how to correct closed loops, and how to check reports for it.
Prerequisites¶
- Python 3.11+,
pip install aiohttp. - An open-loop generator, from building an open-loop load generator in asyncio.
- Percentile basics, from measuring latency percentiles without averaging them.
1. See the omission happen¶
A closed-loop worker measures each request from when it was actually sent. During a stall, it sends one request, waits a second for it, and sends nothing else:
async def worker(session, url, latencies, end):
while time.perf_counter() < end:
t = time.perf_counter()
async with session.get(url) as response:
await response.read()
latencies.append(time.perf_counter() - t) # one 1,000 ms sample per worker per stall
# 10 workers, 6 s, 1 s stall: 4,287 samples, 10 of them slow
# p50 11 ms, p99 13 ms, p99.9 1,011 ms
Ten slow samples out of 4,287 is 0.23%, below the p99 threshold, so the p99 reads 13 ms. A user population arriving at the normal rate during that second would have produced hundreds of slow requests — the generator simply did not send them. The longer the stall relative to the run, and the fewer the workers, the larger the error; at the extreme, a stall of the whole run produces exactly one slow sample per worker.
Verify: compare the request count of a run with a stall to one without; a closed loop sends fewer, and every missing request is an omitted sample.
2. Correct a closed loop with a schedule¶
When you must use a fixed number of workers — a protocol that needs one connection per client, a tool that works that way — give each worker a schedule and measure from the scheduled times:
async def scheduled_worker(session, url, k: int, workers: int, rate: float, seconds: float,
latencies: list[float], start: float) -> None:
interval = workers / rate # each worker's share of the rate
i = 0
while True:
intended = start + k * interval / workers + i * interval
if intended - start > seconds:
return
if (delay := intended - time.perf_counter()) > 0:
await asyncio.sleep(delay)
async with session.get(url) as response: # may start late if the last one was slow
await response.read()
latencies.append(time.perf_counter() - intended)
i += 1
Measured with ten workers at a combined 500 per second: from actual send times the p99 still read 13 ms, but from intended times it read 987 ms. After a stall, each worker's next requests are late, and measuring from when they should have gone out charges the wait to them — the correction used by wrk2 and HdrHistogram. It recovers the latency users would have seen, though not the extra load a real population would have added during the stall; an open loop captures both.
Verify: for the same stall, the corrected closed loop's p99 is close to the open loop's, and far above the uncorrected one.
3. Prefer an open loop for latency questions¶
An open-loop generator never waits for responses before sending, so it neither omits samples nor reduces load during slowdowns:
for i in range(int(rate * seconds)):
intended = start + i / rate
if (delay := intended - time.perf_counter()) > 0:
await asyncio.sleep(delay)
tasks.append(asyncio.create_task(one(intended))) # independent of earlier responses
Measured: 3,000 requests at 500 per second, about 500 of them affected by the stall, and a p99 of 1,472 ms. The open loop also shows the stall's aftermath: the requests that arrived during the stall queued up and were served afterwards, so latency stayed high for a while after the service recovered — real behaviour that a closed loop smooths away. For a question like "what latency do users see at 500 requests per second?", this is the measurement that answers it.
Verify: the open-loop request count is identical with and without the stall.
4. Check existing results for the error¶
Many reported numbers — from tools, papers, colleagues — come from closed loops. A few questions reveal whether coordinated omission affects them:
def suspicious(report: dict) -> list[str]:
warnings = []
if report.get("mode") == "closed" and not report.get("latency_from_intended"):
warnings.append("closed loop measured from actual send times: stalls are under-sampled")
if report["achieved_rps"] < 0.95 * report.get("target_rps", report["achieved_rps"]):
warnings.append("achieved rate below target: requests were not sent as scheduled")
if report["p999_ms"] > 20 * report["p99_ms"]:
warnings.append("p99.9 far above p99: a stall may be hiding in very few samples")
return warnings
The measured closed-loop run tripped the last check: p99.9 (1,011 ms) was 78 times the p99 (13 ms). A large gap between p99 and the maximum, together with a lower request count than expected, is the signature of omitted samples. Tools that implement the correction say so (wrk2's -R rate option, HdrHistogram's expected-interval recording); classic wrk and simple worker pools do not.
Verify: each latency number you rely on states whether it was measured open-loop or with intended-time correction.
5. Apply the same idea to production metrics¶
Coordinated omission is not only a load-testing problem. Any measurement that is taken only when something completes can under-sample slow periods:
# Under-samples stalls: one observation per completed loop iteration
while True:
t = time.perf_counter()
await process_next_batch()
BATCH_LATENCY.observe(time.perf_counter() - t)
# Samples on a schedule: a stalled batch is observed every second it is stalled
async def in_progress_monitor(state):
while True:
await asyncio.sleep(1)
if state.started is not None:
IN_PROGRESS_AGE.set(time.perf_counter() - state.started)
Production latency histograms of requests are mostly fine — each real request is recorded — but internal loops, health checks and synthetic probes that run one at a time have the same flaw as a closed-loop worker. Report the age of in-progress work on a schedule as well, and event-loop lag from a periodic heartbeat, as in alerting on event loop saturation, so stalls are visible while they last.
Verify: a deliberate one-minute stall of a background loop shows up as a rising in-progress age every second of that minute.
Verification¶
Latency benchmarks avoid coordinated omission when:
- Load is open-loop, or closed loops measure from scheduled send times.
- Achieved rate matches target, and request counts do not drop during slowdowns.
- Reports state their method, and large p99.9/p99 gaps are investigated.
- Production probes and loops sample on a schedule, not only on completion.
Diagnostic Hook: plot request count per second alongside latency for every load-test run. A dip in request count that lines up with a latency spike is coordinated omission in action — the generator stopped sending exactly when the service was slow.
Pitfalls & edge cases¶
- Trusting closed-loop p99s. Measured: 13 ms reported through a 1 s stall.
- Few workers. Fewer workers mean fewer samples of any stall.
- Correcting latency but not load. A corrected closed loop still sends less during stalls.
- Completion-only metrics. Stalled loops report nothing while stalled.
Frequently Asked Questions¶
What is coordinated omission in load testing?
The error where a generator waits for slow responses and therefore does not send, or record, the requests that would have experienced the slowdown. In testing, it turned a 1 s stall into a p99 of 13 ms.
How do I avoid coordinated omission?
Use an open-loop generator that sends on a schedule regardless of responses, or give closed-loop workers a schedule and measure latency from each request's intended send time.
Does wrk suffer from coordinated omission?
Classic wrk is a closed-loop tool and measures from actual send times; wrk2 adds a target rate and intended-time correction to address it.
How can I tell if a benchmark report is affected?
Check whether the generator was closed-loop, whether the achieved rate fell below the target, and whether p99.9 or max is vastly larger than p99; together they indicate omitted samples.
Related¶
- Load Testing & Benchmarking — up to the topic overview.
- Finding the saturation point of an async service — where measuring correctly matters most.
- Resilience, Cancellation & Error Handling — the section overview.