Skip to content

Deriving Timeouts from Latency Percentiles

A timeout is a bet about how long a healthy call can take. Too long and a hung dependency holds requests, connections and memory for the full duration; too short and healthy calls fail. The way to choose is from the dependency's measured latency distribution — and the shape of that distribution matters more than any single number. Simulated on 200,000 calls drawn from a realistic distribution — a lognormal body with a median of 40 ms and a 0.5% slow tail of 0.3–1.5 s (GC pauses, cache misses): p90 was 63 ms, p99 98 ms, and p99.9 jumped to 1,262 ms, inside the tail. A timeout at p99 failed 1.00% of healthy calls; at twice p99, 0.50% — exactly the slow tail; at p99.9, 0.10%. Combining a p99 timeout with one retry for idempotent calls gave an end-to-end p99.9 of 160 ms and 0.007% failures, against 1,188 ms and 0.095% for simply waiting up to p99.9. This guide derives timeouts from data and decides when retrying beats waiting.

Prerequisites

1. Measure the distribution per dependency and operation

Timeouts belong to operations, not to clients: a cache lookup and a report query on the same database have nothing in common. Record latency per dependency and operation with enough tail resolution:

from prometheus_client import Histogram

DEP_LATENCY = Histogram(
    "dependency_latency_seconds", "Latency of calls to dependencies",
    ["dependency", "operation", "outcome"],
    buckets=(0.005, 0.01, 0.025, 0.05, 0.075, 0.1, 0.15, 0.25, 0.5, 1, 2.5, 5, 10),
)


async def timed(dependency: str, operation: str, coro):
    start = time.perf_counter()
    outcome = "ok"
    try:
        return await coro
    except Exception:
        outcome = "error"
        raise
    finally:
        DEP_LATENCY.labels(dependency, operation, outcome).observe(time.perf_counter() - start)

Bucket boundaries decide how precisely you can read percentiles later; put several buckets around the expected p99. Use a few weeks of data, including busy periods, and exclude incidents — the timeout describes healthy behaviour.

Verify: for each dependency operation you can read p50, p99 and p99.9 for the last two weeks.

Latency percentiles of a dependency with a slow tail 4 horizontal bars comparing p50 with the others. Latency percentiles of a dependency with a slow tail p50 40 ms p90 63 ms p99 98 ms p99.9 1,262 ms Lognormal body (median 40 ms) plus a 0.5% tail of 0.3-1.5 s; 200,000 simulated calls. Between p99 and p99.9 the distribution changes regime; that is where to look.

2. Read the tail's shape before picking a percentile

The percentiles above show two populations: most calls are fast, and a small share — here 0.5% — are slow for a different reason. A timeout's effect depends on which side of that gap it falls:

def false_timeout_rate(latencies: list[float], timeout: float) -> float:
    return sum(1 for x in latencies if x > timeout) / len(latencies)

# measured on the simulated distribution:
#   timeout = p99   (98 ms)    -> 1.00% of healthy calls time out
#   timeout = 2 x p99 (196 ms) -> 0.50%   (exactly the slow tail: the body is fully covered)
#   timeout = p99.9 (1262 ms)  -> 0.10%
#   timeout = 5 s (default)    -> 0.00%, but a hang holds the request for 5 s

A timeout just above the fast population (196 ms here) cuts off the slow tail and nothing else: the tail calls would have taken 300–1,500 ms anyway. Waiting for them (p99.9) avoids failing them but makes every tail call slow for the caller. Defaults like 5 s or 30 s come from libraries, not from your dependency, and in an outage they convert every request into a 5-second request.

Verify: plot the latency histogram on a log scale; identify whether there is a gap between the main body and a slow tail.

3. Prefer a short timeout plus one retry for idempotent calls

When the slow tail is caused by something transient — a GC pause, an unlucky replica, a cold cache — a second attempt usually lands in the fast population. Simulated, with each retry an independent draw:

async def call_with_retry_on_timeout(fn, per_attempt: float, attempts: int = 2):
    for attempt in range(attempts):
        try:
            async with asyncio.timeout(per_attempt):
                return await fn()
        except TimeoutError:
            if attempt == attempts - 1:
                raise

# p99 timeout (98 ms) + 1 retry:  end-to-end p99.9 = 160 ms, failures 0.007%
# p99.9 timeout (1,262 ms), no retry: end-to-end p99.9 = 1,188 ms, failures 0.095%

The retry costs at most one extra p99-length wait for the 1% that time out, and adds about 1% more load on the dependency; in exchange the tail shrinks from over a second to 160 ms and failures fall more than tenfold. This only works for idempotent operations and only when slow calls are independent — if the slowness is shared (the database is overloaded), retries add load to the problem, which is why they need a budget, as in implementing retry budgets to prevent retry storms. Hedged requests are the parallel version of the same idea, covered in hedging requests to cut tail latency.

Verify: after switching, the dependency's request rate rises by about the timeout rate, and end-to-end tail latency falls.

Timeout strategies on the same simulated dependency A grid of 2 rows by 4 columns. Timeout strategies on the same simulated dependency strategy end-to-end p99 end-to-end p99.9 failures timeout p99 (98 ms) + 1 retry 97 ms 160 ms 0.007% timeout p99.9 (1,262 ms), no retry 98 ms 1,188 ms 0.095% For idempotent calls with independent slow tails, retrying beats waiting.

4. Fit the timeout inside the caller's budget

A dependency timeout must also fit the request that makes the call. The derived value is an upper bound; the remaining request budget may be lower:

DERIVED = {("payments", "authorize"): 0.25, ("search", "query"): 0.15, ("reports", "build"): 8.0}


async def call(dependency: str, operation: str, fn):
    derived = DERIVED[(dependency, operation)]
    left = remaining()                                  # from the request's deadline
    timeout = derived if left is None else min(derived, left)
    async with asyncio.timeout(timeout):
        return await timed(dependency, operation, fn())

Taking the minimum of the derived timeout and the request's remaining time keeps deep call chains from exceeding the outer deadline. Keep derived values in configuration, with the percentile they came from and the date, so they can be re-derived when the dependency changes; latency distributions drift with data size, traffic and deployments. The request-level budget comes from the edge, as in enforcing request timeouts in ASGI servers.

Verify: no dependency call can outlast its request's deadline, checked by a test with a deadline shorter than the derived timeout.

5. Re-derive regularly and alert on drift

Timeouts set once go stale. Compute the false-timeout rate continuously and compare it with what you designed for:

# Alert when healthy-period timeouts exceed the design rate by a margin
- alert: DependencyTimeoutRateHigh
  expr: |
    sum by (dependency, operation) (rate(dependency_timeouts_total[30m]))
      / sum by (dependency, operation) (rate(dependency_latency_seconds_count[30m])) > 0.02
  for: 30m
  annotations:
    summary: "{{ $labels.dependency }}/{{ $labels.operation }} times out on >2% of calls"

A timeout rate that creeps upward without an incident means the dependency's latency has drifted past the derived value — re-derive, or investigate the dependency. A rate that drops to zero for months may mean the timeout is far too generous for current latency and would hold requests needlessly during an outage. Review timeouts on a schedule (quarterly, or with each major dependency change).

Verify: each timeout's configured value is within a factor of two of its current measured percentile.

How should this call's timeout be set? A decision on What does the latency distribution look like with 4 outcomes. How should this call's timeout be set? What does the latency distribution look like? fast body, then a slow tail just above the body cuts only the tail idempotent, independent tail ~p99 + 1 retry p99.9 1,188 -> 160 ms not idempotent nearer p99.9, no retry fewer false timeouts always min(derived, request remaining) re-derive on drift Derive from data, cap by the caller, revisit as data changes.

Verification

Timeouts are well derived when:

  • Each dependency operation has its own value, from measured percentiles.
  • The distribution's shape (body and tail) informed the choice, not one percentile alone.
  • Idempotent calls use a short timeout plus a budgeted retry where the tail is independent.
  • Values are capped by request deadlines and re-derived when latency drifts.

Diagnostic Hook: for each dependency operation, chart the configured timeout as a line over its latency percentiles. A timeout far above p99.9 is not protecting anything during outages; one below p99 is failing healthy calls; one sitting just above the main body of a bimodal distribution is doing exactly its job.

Pitfalls & edge cases

  • Library defaults. A 5 s default turns an outage into 5-second requests.
  • One timeout per client. Operations on the same dependency differ by orders of magnitude.
  • Retrying correlated slowness. If the dependency is overloaded, retries add load.
  • Never re-deriving. Latency drifts; timeouts must follow.

Frequently Asked Questions

How do I choose a timeout value for a dependency call?

From that operation's measured latency: look at p99 and p99.9 and the shape between them, and set the timeout just above the healthy body. In simulation, a timeout at twice p99 failed only the 0.5% slow tail.

Is it better to wait longer or retry after a short timeout?

For idempotent calls whose slow cases are independent, retry: a p99 timeout plus one retry gave an end-to-end p99.9 of 160 ms against 1,188 ms for waiting up to p99.9, with fewer failures.

Should timeouts be the same for all calls to a service?

No. Set them per operation, since latency differs widely between, say, a cache lookup and a report query on the same database.

How often should timeouts be revisited?

Whenever a dependency changes significantly and on a regular schedule; alert when the timeout rate during healthy periods drifts above the designed rate.