Skip to content

Tuning Circuit Breaker Thresholds from Metrics

A circuit breaker's thresholds decide two things that pull in opposite directions: how quickly it stops sending traffic to a dependency that has failed, and how often it trips on a dependency that is merely noisy. Guessing them produces breakers that never open or that flap all day. Simulated on one 100-second trace at 100 requests per second — 2% background errors, a 10-second outage, then 20 seconds at 20% errors — five configurations behaved very differently. Five consecutive failures opened 0.04 s into the outage and sent only 6 requests to the dead dependency, but also tripped once during the 20% period and rejected 499 requests that would mostly have succeeded. Three in a row tripped three times there and rejected 1,372. A 50% failure rate over 10 s never tripped falsely but took 4.86 s to open and sent 488 requests into the outage. 50% over 2 s opened in 0.98 s, sent 100, and never tripped falsely. This guide derives thresholds from your own metrics instead of guessing.

Prerequisites

1. Measure the dependency's normal error rate and volume

Thresholds are only meaningful relative to what normal looks like. Pull two numbers per dependency from a few weeks of metrics: the error rate during healthy periods, and the request rate at quiet times:

# PromQL, per dependency:
#   healthy error rate (p99 over 5-minute windows, excluding incidents)
#     quantile_over_time(0.99, (sum(rate(dep_errors_total[5m])) / sum(rate(dep_requests_total[5m])))[14d:5m])
#   quiet-time volume
#     quantile_over_time(0.05, sum(rate(dep_requests_total[1m]))[14d:1m])

baseline_error_rate = 0.02     # 2% background errors in the simulation
quiet_rps = 100

The background error rate sets how low the threshold may go without false trips; the quiet-time volume sets how short the window may be while still containing enough requests for a rate to mean anything. A dependency at 2 requests per second needs a much longer window than one at 2,000. Define "error" deliberately: timeouts and 5xx count; 4xx caused by the caller's own bad input should not open a breaker that protects everyone.

Verify: you can state, per dependency, its normal error rate and its minimum request rate.

2. Prefer a failure rate over a window to consecutive counts

Consecutive-failure breakers are simple and fast, and their false-trip rate depends entirely on the error rate. With independent errors at rate p, n failures in a row happen about once every 1/pⁿ requests — rare at 2%, frequent at 20%:

from collections import deque


class RateWindow:
    def __init__(self, window: float, min_requests: int, threshold: float) -> None:
        self.window, self.min_requests, self.threshold = window, min_requests, threshold
        self.events: deque[tuple[float, bool]] = deque()

    def record(self, now: float, ok: bool) -> bool:
        """Record a result; return True if the breaker should open."""
        self.events.append((now, ok))
        while self.events and self.events[0][0] < now - self.window:
            self.events.popleft()
        if len(self.events) < self.min_requests:
            return False                              # not enough evidence yet
        failures = sum(1 for _, good in self.events if not good)
        return failures / len(self.events) >= self.threshold

Simulated: during the 20% error period, 5-in-a-row tripped once and 3-in-a-row three times, each trip rejecting traffic for the open period — 499 and 1,372 rejected requests, most of which would have succeeded. The 50% rate thresholds never tripped there, because 20% is far from 50%. A rate threshold separates "degraded" from "down"; a consecutive count cannot.

Verify: replay a recorded day of request outcomes through the candidate breaker offline; it should not open outside real incidents.

Five breaker configurations on the same 100 s trace A grid of 5 rows by 5 columns. Five breaker configurations on the same 100 s trace configuration opens after sent in outage trips at 20% errors rejected at 20% 5 consecutive failures 0.04 s 6 1 499 3 consecutive failures 0.02 s 4 3 1,372 rate >= 50% / 10 s, min 20 4.86 s 488 0 0 rate >= 50% / 2 s, min 20 0.98 s 100 0 0 rate >= 10% / 10 s, min 20 0.80 s 82 3 1,502 Virtual-time simulation, 100 requests/s, open period 5 s.

3. Choose the window from detection time

With a rate threshold, the window length controls detection time. In a full outage, the failure fraction in the window climbs from the baseline toward 100%, crossing a 50% threshold roughly halfway through the window:

def time_to_open(window: float, threshold: float, baseline: float) -> float:
    """Approximate seconds from outage start until the windowed rate crosses threshold."""
    return window * (threshold - baseline) / (1 - baseline)


time_to_open(10, 0.5, 0.02)     # ~4.9 s   (simulated: 4.86 s)
time_to_open(2, 0.5, 0.02)      # ~0.98 s  (simulated: 0.98 s)

The estimate matched the simulation exactly. Choose the window from how long you are willing to keep sending traffic into an outage — every second costs the dependency's recovery and your users' latency — and check that the window still holds at least min_requests at quiet-time volume. With 100 requests per second, 2 s holds 200 requests, plenty; at 5 requests per second it holds 10, and the window must grow or the minimum must shrink.

Verify: at quiet-time volume, the chosen window contains at least min_requests requests.

Requests sent into a 10 s outage before the breaker opened 4 horizontal bars comparing 5 consecutive failures with the others. Requests sent into a 10 s outage before the breaker opened 5 consecutive failures 6 rate >= 10% / 10 s 82 rate >= 50% / 2 s 100 rate >= 50% / 10 s 488 Includes half-open probes during the outage. Detection speed is window x threshold; false trips come from the threshold alone.

4. Set the threshold between degraded and down

The threshold should sit well above the worst error rate you are prepared to tolerate as "degraded but useful" and well below "down":

BREAKERS = {
    # name:        (window s, min requests, failure-rate threshold, open s)
    "payments":    (2.0, 20, 0.50, 5.0),     # high volume, critical: fast detection
    "search":      (5.0, 20, 0.50, 10.0),    # moderate volume
    "geo-lookup":  (30.0, 10, 0.60, 30.0),   # low volume: longer window
}

A 10% threshold opened faster in the outage (0.80 s) but tripped three times during the 20% period and rejected 1,502 requests — if 20% errors are an acceptable degraded mode, 10% is too sensitive. When a dependency degrades gradually, a breaker is the wrong tool for the middle ground; timeouts, retries with budgets and fallbacks handle it, as in adding timeouts and fallbacks to degraded dependencies. The breaker is for "this is not going to work right now".

Verify: the threshold sits at least a few times above the healthy error rate, and above the degraded rate you accept.

5. Validate by replaying real traffic

Before changing production thresholds, replay recorded outcomes — including a past incident — through the candidate configuration and count what it would have done:

def replay(events: list[tuple[float, bool]], breaker_factory) -> dict:
    breaker = breaker_factory()
    stats = {"sent": 0, "rejected": 0, "opens": 0}
    for ts, ok in events:
        if breaker.allow(ts):
            stats["sent"] += 1
            was_open = breaker.state == "open"
            breaker.record(ts, ok)
            stats["opens"] += breaker.state == "open" and not was_open
        else:
            stats["rejected"] += 1
    return stats

The simulation in this guide is exactly such a replay on a synthetic trace; the same code on a day of real outcomes, exported from request logs, tells you how many false trips a threshold would have caused and how much load it would have kept off a failing dependency. Re-run it when traffic patterns change, and alert on breaker state transitions so each trip in production is reviewed.

Verify: the replay of a known incident opens the breaker within your detection target, and a replay of a normal week opens it zero times.

Which breaker trigger fits this dependency? A decision on What do the dependency's metrics show with 4 outcomes. Which breaker trigger fits this dependency? What do the dependency's metrics show? high volume, some background errors rate >= 50% over ~2 s fast, no false trips low volume longer window, lower minimum enough samples near-zero healthy errors consecutive failures ok simplest degraded mode is acceptable threshold above it fallbacks handle the middle Window sets detection time; threshold sets false-trip rate.

Verification

Breaker thresholds are tuned when:

  • They are derived from measured baseline error rates and volumes.
  • A failure rate over a window is used wherever errors are not near zero.
  • The window gives the detection time you want and holds enough requests at quiet times.
  • Replays of real traffic show no false trips and fast opening in incidents.

Diagnostic Hook: export breaker state transitions with the error rate and request count in the window at the moment of opening. Opens with a window error rate near the threshold during normal operation are false trips — raise the threshold or lengthen the window; opens that come long after an incident started mean the window is too long.

Pitfalls & edge cases

  • Consecutive-failure breakers on noisy dependencies. Simulated: 1,372 needless rejections at 20% errors.
  • Long windows. Simulated: 4.86 s and 488 requests into a dead dependency.
  • No minimum request count. Two failures out of three trip a rate breaker at quiet times.
  • Counting caller errors. 4xx from bad input should not open a shared breaker.

Frequently Asked Questions

What failure threshold should a circuit breaker use?

Set it well above the dependency's normal and acceptable degraded error rates and well below total failure; 50% is a common choice. In simulation, 50% never tripped on a 20%-error period, while 10% tripped three times.

Should a circuit breaker count consecutive failures or a failure rate?

A failure rate over a time window with a minimum request count, unless the dependency's healthy error rate is near zero. Consecutive counts tripped repeatedly at 20% errors in simulation.

How long should the circuit breaker window be?

Long enough to hold a meaningful number of requests at quiet times, short enough to detect outages quickly. Opening time is roughly window x threshold; a 2 s window at 50% opened in 0.98 s.

How do I test circuit breaker settings safely?

Replay recorded request outcomes, including a past incident, through the candidate configuration offline and count trips, rejections and requests sent during the outage.