Tuning Circuit Breaker Thresholds from Metrics¶
A circuit breaker's thresholds decide two things that pull in opposite directions: how quickly it stops sending traffic to a dependency that has failed, and how often it trips on a dependency that is merely noisy. Guessing them produces breakers that never open or that flap all day. Simulated on one 100-second trace at 100 requests per second — 2% background errors, a 10-second outage, then 20 seconds at 20% errors — five configurations behaved very differently. Five consecutive failures opened 0.04 s into the outage and sent only 6 requests to the dead dependency, but also tripped once during the 20% period and rejected 499 requests that would mostly have succeeded. Three in a row tripped three times there and rejected 1,372. A 50% failure rate over 10 s never tripped falsely but took 4.86 s to open and sent 488 requests into the outage. 50% over 2 s opened in 0.98 s, sent 100, and never tripped falsely. This guide derives thresholds from your own metrics instead of guessing.
Prerequisites¶
- Python 3.11+; the simulation and breaker code are stdlib only.
- A breaker implementation, from implementing an async circuit breaker.
- Per-dependency metrics: request rate, error rate and latency, from exporting Prometheus metrics from asyncio.
1. Measure the dependency's normal error rate and volume¶
Thresholds are only meaningful relative to what normal looks like. Pull two numbers per dependency from a few weeks of metrics: the error rate during healthy periods, and the request rate at quiet times:
# PromQL, per dependency:
# healthy error rate (p99 over 5-minute windows, excluding incidents)
# quantile_over_time(0.99, (sum(rate(dep_errors_total[5m])) / sum(rate(dep_requests_total[5m])))[14d:5m])
# quiet-time volume
# quantile_over_time(0.05, sum(rate(dep_requests_total[1m]))[14d:1m])
baseline_error_rate = 0.02 # 2% background errors in the simulation
quiet_rps = 100
The background error rate sets how low the threshold may go without false trips; the quiet-time volume sets how short the window may be while still containing enough requests for a rate to mean anything. A dependency at 2 requests per second needs a much longer window than one at 2,000. Define "error" deliberately: timeouts and 5xx count; 4xx caused by the caller's own bad input should not open a breaker that protects everyone.
Verify: you can state, per dependency, its normal error rate and its minimum request rate.
2. Prefer a failure rate over a window to consecutive counts¶
Consecutive-failure breakers are simple and fast, and their false-trip rate depends entirely on the error rate. With independent errors at rate p, n failures in a row happen about once every 1/pⁿ requests — rare at 2%, frequent at 20%:
from collections import deque
class RateWindow:
def __init__(self, window: float, min_requests: int, threshold: float) -> None:
self.window, self.min_requests, self.threshold = window, min_requests, threshold
self.events: deque[tuple[float, bool]] = deque()
def record(self, now: float, ok: bool) -> bool:
"""Record a result; return True if the breaker should open."""
self.events.append((now, ok))
while self.events and self.events[0][0] < now - self.window:
self.events.popleft()
if len(self.events) < self.min_requests:
return False # not enough evidence yet
failures = sum(1 for _, good in self.events if not good)
return failures / len(self.events) >= self.threshold
Simulated: during the 20% error period, 5-in-a-row tripped once and 3-in-a-row three times, each trip rejecting traffic for the open period — 499 and 1,372 rejected requests, most of which would have succeeded. The 50% rate thresholds never tripped there, because 20% is far from 50%. A rate threshold separates "degraded" from "down"; a consecutive count cannot.
Verify: replay a recorded day of request outcomes through the candidate breaker offline; it should not open outside real incidents.
3. Choose the window from detection time¶
With a rate threshold, the window length controls detection time. In a full outage, the failure fraction in the window climbs from the baseline toward 100%, crossing a 50% threshold roughly halfway through the window:
def time_to_open(window: float, threshold: float, baseline: float) -> float:
"""Approximate seconds from outage start until the windowed rate crosses threshold."""
return window * (threshold - baseline) / (1 - baseline)
time_to_open(10, 0.5, 0.02) # ~4.9 s (simulated: 4.86 s)
time_to_open(2, 0.5, 0.02) # ~0.98 s (simulated: 0.98 s)
The estimate matched the simulation exactly. Choose the window from how long you are willing to keep sending traffic into an outage — every second costs the dependency's recovery and your users' latency — and check that the window still holds at least min_requests at quiet-time volume. With 100 requests per second, 2 s holds 200 requests, plenty; at 5 requests per second it holds 10, and the window must grow or the minimum must shrink.
Verify: at quiet-time volume, the chosen window contains at least min_requests requests.
4. Set the threshold between degraded and down¶
The threshold should sit well above the worst error rate you are prepared to tolerate as "degraded but useful" and well below "down":
BREAKERS = {
# name: (window s, min requests, failure-rate threshold, open s)
"payments": (2.0, 20, 0.50, 5.0), # high volume, critical: fast detection
"search": (5.0, 20, 0.50, 10.0), # moderate volume
"geo-lookup": (30.0, 10, 0.60, 30.0), # low volume: longer window
}
A 10% threshold opened faster in the outage (0.80 s) but tripped three times during the 20% period and rejected 1,502 requests — if 20% errors are an acceptable degraded mode, 10% is too sensitive. When a dependency degrades gradually, a breaker is the wrong tool for the middle ground; timeouts, retries with budgets and fallbacks handle it, as in adding timeouts and fallbacks to degraded dependencies. The breaker is for "this is not going to work right now".
Verify: the threshold sits at least a few times above the healthy error rate, and above the degraded rate you accept.
5. Validate by replaying real traffic¶
Before changing production thresholds, replay recorded outcomes — including a past incident — through the candidate configuration and count what it would have done:
def replay(events: list[tuple[float, bool]], breaker_factory) -> dict:
breaker = breaker_factory()
stats = {"sent": 0, "rejected": 0, "opens": 0}
for ts, ok in events:
if breaker.allow(ts):
stats["sent"] += 1
was_open = breaker.state == "open"
breaker.record(ts, ok)
stats["opens"] += breaker.state == "open" and not was_open
else:
stats["rejected"] += 1
return stats
The simulation in this guide is exactly such a replay on a synthetic trace; the same code on a day of real outcomes, exported from request logs, tells you how many false trips a threshold would have caused and how much load it would have kept off a failing dependency. Re-run it when traffic patterns change, and alert on breaker state transitions so each trip in production is reviewed.
Verify: the replay of a known incident opens the breaker within your detection target, and a replay of a normal week opens it zero times.
Verification¶
Breaker thresholds are tuned when:
- They are derived from measured baseline error rates and volumes.
- A failure rate over a window is used wherever errors are not near zero.
- The window gives the detection time you want and holds enough requests at quiet times.
- Replays of real traffic show no false trips and fast opening in incidents.
Diagnostic Hook: export breaker state transitions with the error rate and request count in the window at the moment of opening. Opens with a window error rate near the threshold during normal operation are false trips — raise the threshold or lengthen the window; opens that come long after an incident started mean the window is too long.
Pitfalls & edge cases¶
- Consecutive-failure breakers on noisy dependencies. Simulated: 1,372 needless rejections at 20% errors.
- Long windows. Simulated: 4.86 s and 488 requests into a dead dependency.
- No minimum request count. Two failures out of three trip a rate breaker at quiet times.
- Counting caller errors. 4xx from bad input should not open a shared breaker.
Frequently Asked Questions¶
What failure threshold should a circuit breaker use?
Set it well above the dependency's normal and acceptable degraded error rates and well below total failure; 50% is a common choice. In simulation, 50% never tripped on a 20%-error period, while 10% tripped three times.
Should a circuit breaker count consecutive failures or a failure rate?
A failure rate over a time window with a minimum request count, unless the dependency's healthy error rate is near zero. Consecutive counts tripped repeatedly at 20% errors in simulation.
How long should the circuit breaker window be?
Long enough to hold a meaningful number of requests at quiet times, short enough to detect outages quickly. Opening time is roughly window x threshold; a 2 s window at 50% opened in 0.98 s.
How do I test circuit breaker settings safely?
Replay recorded request outcomes, including a past incident, through the candidate configuration offline and count trips, rejections and requests sent during the outage.
Related¶
- Circuit Breakers & Bulkheads — up to the topic overview.
- Probing strategies for half-open circuit breakers — what happens after the breaker opens.
- Resilience, Cancellation & Error Handling — the section overview.