Per-Endpoint Circuit Breakers¶
A circuit breaker protects one failure domain. Choosing what it is keyed by — the host, the endpoint, the full URL — decides whether a failure in one part of a dependency blocks the healthy parts, and whether the breaker ever opens at all. Measured on Python 3.14 with 20 clients calling a simulated API for 4 seconds, half the calls to /search and half to /profile, while /search failed for 2 seconds and /profile stayed healthy, each breaker opening after 5 failures and probing after 0.5 s: one breaker for the host rejected 3,834 calls to the healthy /profile, and still let 184 /search calls reach the failing endpoint, because successful /profile probes kept closing it. One breaker per endpoint rejected 0 /profile calls and let only 11 failing /search calls through. Keying by the raw URL path, including the item ID, created 10,627 breakers, none of which ever opened, and 2,781 failing calls reached /search. This guide picks the key and builds the registry.
Prerequisites¶
- An async circuit breaker, from implementing an async circuit breaker or a library.
- Library behaviour, from using circuit breaker libraries.
- The topic overview, Circuit Breakers & Bulkheads.
1. Measure what a per-host breaker does¶
One breaker for a whole API is the simplest setup:
breaker = Breaker(threshold=5, reset_after=0.5)
async def call(endpoint, item_id):
return await breaker.call(api, endpoint, item_id)
Measured during the /search outage: the breaker opened on /search failures and then rejected everything, including 3,834 calls to /profile, which would have succeeded — 2,700 /profile calls succeeded over the whole run, so more were rejected than served. Worse, it did not protect /search well: each time the breaker went half-open, the probe was as likely to be a /profile call as a /search one, and a successful /profile probe closed the breaker, sending a fresh wave of calls to the failing endpoint. 184 /search calls reached it and failed. A shared breaker turns a partial outage into a full one and still leaks traffic to the broken part.
Verify: during a test outage of one endpoint, count rejections on the other endpoints; with a per-host breaker they will be high.
2. Key breakers by endpoint template¶
The failure domain of most APIs is the endpoint — a route handled by one backend code path and its own dependencies. Key the breaker by method and route template:
class BreakerRegistry:
def __init__(self, factory):
self._factory = factory
self._breakers: dict[str, Breaker] = {}
def for_key(self, key: str) -> Breaker:
breaker = self._breakers.get(key)
if breaker is None:
breaker = self._breakers[key] = self._factory()
return breaker
registry = BreakerRegistry(lambda: Breaker(threshold=5, reset_after=30))
async def call(method: str, template: str, **params):
breaker = registry.for_key(f"{method} api.example{template}")
return await breaker.call(api, method, template.format(**params))
Measured: with two breakers, one per endpoint, the /search breaker opened and stayed open between failed probes, letting only 11 failing calls through, and the /profile breaker never opened — 0 of its calls were rejected and 5,750 succeeded. Each part of the API was protected from its own failures and only its own.
Verify: an outage injected on one endpoint opens only that endpoint's breaker.
3. Normalize URLs before keying¶
Keys must have a bounded number of values. A key built from the request path includes IDs, query strings and slugs, so every key sees only a handful of calls. Measured: keying by the raw path with an item ID produced 10,627 breakers in 4 seconds; none reached 5 consecutive failures, so none opened, and 2,781 calls reached the failing endpoint — protection equal to no breaker at all, plus an unbounded dictionary. Normalize to a template:
ID = re.compile(r"/(\d+|[0-9a-f]{8}-[0-9a-f-]{27,})(?=/|$)")
def template_for(path: str) -> str:
path = path.split("?", 1)[0]
return ID.sub("/{id}", path) # /items/123/reviews -> /items/{id}/reviews
assert template_for("/items/123/reviews?page=2") == "/items/{id}/reviews"
Better still, use the route template from the client's own code — the string before formatting — rather than reverse-engineering it from paths. Cap the registry size and alert if it grows, as a guard against a missed normalization.
Verify: the number of breakers stays equal to the number of endpoint templates after a load test with random IDs.
4. Combine endpoint breakers with a host-level view¶
Endpoint breakers miss one failure: the whole host going down. Every endpoint will trip on its own, but each needs its own threshold of failures first, so a host outage costs threshold × endpoints failed calls before everything is open. Add a coarse signal that catches host-wide failures — connection errors, not HTTP errors — without letting endpoint errors trip it:
async def call(method, template, **params):
host_breaker = registry.for_key("host api.example")
endpoint_breaker = registry.for_key(f"{method} api.example{template}")
try:
return await host_breaker.call(endpoint_breaker.call, api, method, template.format(**params))
except HTTPStatusError:
raise # endpoint-level failure: host breaker unaffected
This only works if the host breaker counts connection-level failures — ConnectionError, TLS errors, timeouts on connect — and not 5xx responses from one endpoint, or it re-creates the per-host problem from step 1. Configure what each breaker counts explicitly, as discussed in using circuit breaker libraries.
Verify: a host outage opens the host breaker after its threshold, and an endpoint outage does not.
5. Observe breakers by key¶
With several breakers per dependency, dashboards need the key as a label:
def export(registry: BreakerRegistry):
for key, breaker in registry._breakers.items():
breaker_state.labels(key=key).set({"closed": 0, "half_open": 1, "open": 2}[breaker.state])
breaker_rejections.labels(key=key).set(breaker.rejected_total)
A per-endpoint view shows which part of a dependency is failing, which is often the first useful fact in an incident — "search is down, profiles are fine" rather than "the API is down". Keep the number of label values bounded, which the normalization in step 3 guarantees. For thresholds derived from these metrics, see tuning circuit breaker thresholds from metrics.
Verify: the dashboard shows one series per endpoint template, and an injected endpoint outage appears on that series only.
Verification¶
Per-endpoint breakers work when:
- Breakers are keyed by method and endpoint template, not by host or raw URL.
- The number of breakers is bounded by the number of templates.
- A host-level breaker counts only connection failures, so endpoint errors stay local.
- Metrics carry the breaker key, showing which endpoint is failing.
Diagnostic Hook: when a breaker sees thousands of failures in an incident but never opens, count the breakers in its registry. Keying by raw URL created 10,627 breakers in 4 seconds here, none with enough failures to open.
Pitfalls & edge cases¶
- One breaker per host. Measured: 3,834 healthy calls rejected.
- Probes from healthy endpoints closing a shared breaker. Measured: 184 failing calls leaked.
- Keys with IDs. Measured: 10,627 breakers, none opened.
- A host breaker that counts 5xx responses. It recreates the per-host problem.
Frequently Asked Questions¶
Should a circuit breaker be per host or per endpoint?
Per endpoint template. A per-host breaker rejected 3,834 calls to a healthy endpoint during another endpoint's outage; per-endpoint breakers rejected none.
Why does my circuit breaker never open?
Its key may be too fine. Keying by raw URL with IDs created 10,627 breakers in 4 s, none reaching the failure threshold. Normalize to a template.
How do I handle a whole host failing with per-endpoint breakers?
Add a host-level breaker that counts only connection-level failures, so a host outage opens it while endpoint errors do not.
Why does a shared breaker keep closing during an outage?
Its half-open probe can be a call to a healthy endpoint, whose success closes it. That let 184 failing calls through in this test.
Related¶
- Circuit Breakers & Bulkheads — up to the topic overview.
- Testing circuit breakers — the same tests, per key.
- Resilience, Cancellation & Error Handling — the section overview.