Measuring Latency Percentiles Without Averaging Them¶
Percentiles summarize a distribution, and the summary of a combination is not the combination of the summaries. Averaging p99s across instances, across time windows or across endpoints produces a number that describes no request that ever happened. Simulated with four instances — three healthy, one with a slow dependency affecting 10% of its much smaller traffic: the per-instance p99s were 76, 76, 77 and 969 ms; their average was 299 ms; the true p99 of all requests was 81 ms, and the true p99.9 880 ms. Merging the four instances' histogram buckets and taking the quantile gave 86 ms, within the buckets' resolution. Over time, the mean of sixty one-minute p99s was 108 ms against a true hourly p99 of 98 ms, while the worst minute's p99 was 1,984 ms. This guide combines latency data correctly and shows what to report instead of averages.
Prerequisites¶
- Latency histograms, from exporting Prometheus metrics from asyncio.
- A load generator or production metrics producing per-instance data.
- Python 3.11+ for the examples; PromQL for the dashboard queries.
1. Know why averaged percentiles mislead¶
A p99 is the latency below which 99% of a set of requests fall. Averaging four of them weights each instance equally regardless of traffic, and mixes values from different points of different distributions:
per_instance_p99 = [0.076, 0.076, 0.077, 0.969] # measured
naive = sum(per_instance_p99) / 4 # 0.299 s - matches no request population
all_requests = [x for instance in instances for x in instance]
true_p99 = percentile(all_requests, 0.99) # 0.081 s
true_p999 = percentile(all_requests, 0.999) # 0.880 s
The slow instance served about 3% of the traffic, so its slow requests are about 0.3% of all requests — above the p99 threshold, inside the p99.9. The average put one quarter of the weight on that instance and produced 299 ms, overstating the overall p99 nearly fourfold. With other mixes the same mistake understates: averaging hides one bad instance among many good ones. Either way it is not a percentile of anything.
Verify: compute the p99 of the union of a test's raw latencies and compare it with the average of per-instance p99s; they differ unless the instances are identical.
2. Merge histograms, then take quantiles¶
Raw latencies are rarely kept, but histograms merge exactly: add the counts in each bucket, then compute the quantile from the merged counts. Prometheus histograms are designed for this:
# Correct: sum buckets across instances, then the quantile
histogram_quantile(0.99,
sum by (le) (rate(http_request_duration_seconds_bucket{job="api"}[5m])))
# Wrong: quantile per instance, then averaged
avg(histogram_quantile(0.99, rate(http_request_duration_seconds_bucket{job="api"}[5m])))
Measured: merging the four instances' buckets gave a p99 of 86 ms against the true 81 ms — the difference is bucket resolution, which interpolation within a bucket cannot remove. Put several buckets around the latencies you care about; a bucket edge at 100 ms and the next at 250 ms makes any p99 between them a guess. Summaries (pre-computed quantiles per instance) cannot be merged at all; prefer histograms whenever data from several processes must be combined.
Verify: dashboards compute fleet percentiles with sum by (le) inside histogram_quantile, never avg outside it.
3. Do not average percentiles over time either¶
The same error applies to time windows: the mean of per-minute p99s is not the hourly p99. Measured over sixty minutes with one bad minute, the mean of the per-minute p99s was 108 ms, the true hourly p99 98 ms, and the bad minute's own p99 1,984 ms:
# Hourly p99 from the hour's buckets
histogram_quantile(0.99, sum by (le) (increase(http_request_duration_seconds_bucket[1h])))
# Worst 1-minute p99 within the last hour - shows short incidents the hourly value hides
max_over_time(histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket[1m])))[1h:1m])
Both numbers are useful, for different questions. The hourly p99 says how the hour went for users overall; the worst minute's p99 shows that a short incident happened. The average of minute p99s answers neither. For SLOs, count requests over the threshold instead of using percentiles at all: the share of requests slower than 300 ms over 30 days is directly additive across windows and instances.
Verify: SLO compliance is computed from counts of good and bad requests, not from averaged percentiles.
4. Keep per-dimension signals separately¶
Merging is right for the fleet-wide question; it also dilutes problems confined to one instance, endpoint or tenant. The slow instance's p99 was 969 ms while the fleet's was 81 ms. Keep both views:
# Fleet view: what users overall experience
histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket[5m])))
# Outliers: any instance whose own p99 is far above the fleet's
topk(3, histogram_quantile(0.99, sum by (le, instance) (rate(http_request_duration_seconds_bucket[5m]))))
Alert on the fleet value against the SLO, and on outliers relative to the fleet: one instance with a p99 ten times the fleet's is a broken instance — a bad host, a cold cache, a failing dependency on one path — even if the fleet SLO is met. The same applies per endpoint: merging a fast health check with a slow report endpoint produces a p99 that describes neither, so compute percentiles per route.
Verify: a single slow instance in a test environment triggers the outlier alert without breaching the fleet SLO.
5. Record your own latencies as histograms¶
In load tests and benchmarks you control the recording. Use log-spaced histograms that merge exactly across generator processes and runs:
import bisect
import math
class LogHistogram:
def __init__(self, lowest: float = 1e-4, highest: float = 60.0, per_doubling: int = 8) -> None:
steps = int(math.log2(highest / lowest) * per_doubling) + 1
self.bounds = [lowest * 2 ** (i / per_doubling) for i in range(steps)]
self.counts = [0] * (len(self.bounds) + 1)
def record(self, seconds: float) -> None:
self.counts[bisect.bisect_left(self.bounds, seconds)] += 1
def merge(self, other: "LogHistogram") -> None:
self.counts = [a + b for a, b in zip(self.counts, other.counts)]
def quantile(self, q: float) -> float:
target, seen = q * sum(self.counts), 0
for i, count in enumerate(self.counts):
seen += count
if seen >= target:
return self.bounds[min(i, len(self.bounds) - 1)]
return self.bounds[-1]
Eight buckets per doubling gives about 9% relative precision at every scale, from 0.1 ms to a minute, in 155 counters; HdrHistogram offers the same idea with configurable precision. Each generator process records into its own histogram, and the report merges them — the same discipline as on the dashboard. Keeping the histograms, not only the summary percentiles, lets later analysis ask new questions of old runs, as described in building an open-loop load generator in asyncio.
Verify: merging histograms from two halves of a run gives the same quantiles as recording the whole run in one histogram.
Verification¶
Latency percentiles are handled correctly when:
- No dashboard or report averages percentiles across instances or time.
- Fleet percentiles come from merged histogram buckets, with buckets dense around targets.
- SLOs use counts over a threshold, and outliers are tracked per dimension.
- Load tests record mergeable histograms, not just summary numbers.
Diagnostic Hook: search dashboards and alert rules for avg(histogram_quantile( and avg_over_time( applied to quantiles. Each match is a percentile being averaged; replace it with a quantile of summed buckets, or with a count-based SLO expression.
Pitfalls & edge cases¶
- Averaging per-instance p99s. Simulated: 299 ms reported for a true 81 ms.
- Averaging per-minute p99s. It neither matches the hour nor shows the bad minute.
- Coarse buckets. The merged quantile is only as precise as the bucket edges.
- Merging dissimilar routes. A p99 of health checks and reports describes neither.
Frequently Asked Questions¶
Can you average p99 latency across servers?
No. Averaging percentiles weights servers equally and mixes different distributions; in simulation it gave 299 ms when the true p99 of all requests was 81 ms. Merge histogram buckets and compute the quantile from the merged counts.
How do I compute p99 across instances in Prometheus?
Use histogram_quantile(0.99, sum by (le) (rate(metric_bucket[5m]))), summing buckets across instances before taking the quantile, rather than averaging per-instance quantiles.
Is the average of hourly p99 values the daily p99?
No. Compute the daily p99 from the day's merged buckets, and use the worst window's p99 to find short incidents; the average of window percentiles answers neither question.
How should latency SLOs be measured?
As the share of requests slower than a threshold, counted over the SLO window; counts add correctly across instances and time, unlike percentiles.
Related¶
- Load Testing & Benchmarking — up to the topic overview.
- Deriving timeouts from latency percentiles — percentiles put to work.
- Resilience, Cancellation & Error Handling — the section overview.