Probing Strategies for Half-Open Circuit Breakers¶
When a circuit breaker's open period ends, it moves to half-open and lets some traffic through to test whether the dependency has recovered. The common textbook version allows a single probe and closes completely on success. That works for a dependency that comes back at full strength, and it is exactly wrong for one that comes back fragile — a database still warming its caches, a service that just restarted with cold connection pools, a provider recovering from overload. Simulated with 100 requests per second against a dependency that, for ten seconds after an outage, could handle only 30 per second: closing on the first successful probe sent full traffic at once, caused 118 errors after recovery and 2 re-trips; ramping admitted traffic through 10%, 25%, 50% and 100%, two seconds per step, caused 36 errors and 1 re-trip — at the cost of reaching full traffic 17 s after recovery instead of 13 s. This guide implements both, and chooses between them.
Prerequisites¶
- Python 3.11+, stdlib only; results from a virtual-time simulation.
- A breaker implementation, from implementing an async circuit breaker.
- Threshold tuning, from tuning circuit breaker thresholds from metrics.
1. Understand what half-open is for¶
Half-open answers one question — "is the dependency usable again?" — while risking as little as possible:
class Breaker:
def __init__(self, open_for: float) -> None:
self.state = "closed"
self.open_for = open_for
self.reopen_at = 0.0
def admit(self, now: float) -> bool:
if self.state == "open":
if now < self.reopen_at:
return False # fail fast
self.state = "half_open" # time to test
if self.state == "half_open":
return self.half_open_admit(now) # the strategy lives here
return True
The open period gives the dependency time to recover without your traffic; the half-open policy decides how much traffic tests it, and how quickly to return to normal. Too little probing keeps users on the fallback longer than necessary; too much turns your recovering traffic into the next outage. The right amount depends on how your dependencies recover, which is worth observing in past incidents.
Verify: the breaker exposes its state as a metric, and half-open periods are visible on dashboards alongside the dependency's error rate.
2. Use a single probe for dependencies that recover cleanly¶
Admit one request at a time while half-open; close on success, reopen on failure:
class SingleProbe(Breaker):
def __init__(self, open_for: float) -> None:
super().__init__(open_for)
self.probe_in_flight = False
def half_open_admit(self, now: float) -> bool:
if self.probe_in_flight:
return False # everyone else still fails fast
self.probe_in_flight = True
return True
def record(self, now: float, ok: bool) -> None:
if self.state == "half_open":
self.probe_in_flight = False
if ok:
self.state = "closed" # full traffic immediately
else:
self.state, self.reopen_at = "open", now + self.open_for
Allowing exactly one probe matters in async code: without the probe_in_flight flag, every request that arrives while the first probe is waiting also becomes a probe. Simulated against a fragile recovery, the single probe succeeded as soon as the dependency answered at all, the breaker closed, 100 requests per second hit a dependency that could take 30, errors pushed the window over threshold, and the breaker re-opened — twice — for 118 errors in total. Against a dependency that recovers at full capacity it is ideal: one request of risk, then immediate full service.
Verify: while half-open, exactly one request reaches the dependency at a time.
3. Ramp traffic for dependencies that recover fragile¶
Admit a growing fraction of traffic in steps, checking the error rate at each step before moving on:
import random
class RampingProbe(Breaker):
STEPS = (0.10, 0.25, 0.50, 1.00)
def __init__(self, open_for: float, step_seconds: float = 2.0, max_error_rate: float = 0.5) -> None:
super().__init__(open_for)
self.step_seconds, self.max_error_rate = step_seconds, max_error_rate
self.step = 0
self.step_started = 0.0
self.ok = self.total = 0
def half_open_admit(self, now: float) -> bool:
if self.total == 0 and self.ok == 0 and self.step_started == 0.0:
self.step_started = now
return random.random() < self.STEPS[self.step]
def record(self, now: float, ok: bool) -> None:
if self.state != "half_open":
return
self.total += 1
self.ok += ok
if self.total >= 10 and (self.total - self.ok) / self.total >= self.max_error_rate:
self.state, self.reopen_at = "open", now + self.open_for # back off again
self._reset()
elif now - self.step_started >= self.step_seconds:
if self.step == len(self.STEPS) - 1:
self.state = "closed"
self._reset()
else:
self.step += 1
self.step_started, self.ok, self.total = now, 0, 0
def _reset(self) -> None:
self.step, self.step_started, self.ok, self.total = 0, 0.0, 0, 0
Simulated: the ramp held load at 10 and 25 requests per second during the fragile period — inside the dependency's 30 — re-tripped once when it stepped to 50%, and caused 36 errors against 118. Full traffic resumed after 17 s instead of 13 s, because each step takes time. Rejected requests during the ramp still get the fallback path, so users see degraded service rather than errors.
Verify: during a staged recovery in a test, admitted traffic follows the steps and the dependency's error rate stays under the step threshold.
4. Probe with real requests or a health check¶
Probes can be ordinary user requests, as above, or dedicated health checks sent by the breaker itself:
async def health_probe_loop(breaker: Breaker, client, url: str, interval: float = 1.0) -> None:
while True:
await asyncio.sleep(interval)
if breaker.state != "open" or time.monotonic() < breaker.reopen_at:
continue
try:
async with asyncio.timeout(2.0):
response = await client.get(url)
healthy = response.status_code == 200
except (TimeoutError, httpx.HTTPError):
healthy = False
if healthy:
breaker.state = "half_open" # now let real traffic ramp in
else:
breaker.reopen_at = time.monotonic() + breaker.open_for
Health-check probes cost nothing for users — no real request is risked — but a health endpoint can report healthy while the real operation still fails (a database reachable but a table locked, a service up but its own dependency down). The combination works well: use a health check to leave the open state, then ramp real traffic to confirm. With several processes, each has its own breaker and sends its own probes; share state as in sharing circuit breaker state across processes, or accept N probes for N processes.
Verify: while the dependency is down, probe traffic is bounded to one request per interval per process.
5. Back off the open period on repeated failures¶
A fixed open period re-tests a long outage every few seconds forever. Grow the open period while probes keep failing, and reset it on recovery:
class BackoffOpen:
def __init__(self, base: float = 2.0, cap: float = 60.0) -> None:
self.base, self.cap, self.failures = base, cap, 0
def next_open_for(self) -> float:
self.failures += 1
delay = min(self.cap, self.base * 2 ** (self.failures - 1))
return delay * random.uniform(0.8, 1.2) # jitter across processes
def reset(self) -> None:
self.failures = 0
Short initial open periods recover quickly from brief blips; growing ones stop a fleet of processes from probing a dead dependency every two seconds for an hour. Jitter spreads the probes from many processes so they do not arrive together. The same exponential-backoff reasoning applies as in exponential backoff with jitter in asyncio.
Verify: during a long outage in a test, the interval between probes grows to the cap; after recovery, the next trip starts again at the base.
Verification¶
Half-open behaviour is sound when:
- Only one probe is in flight at a time for single-probe breakers.
- Fragile dependencies are ramped, with an error check at each step.
- Probing load during long outages is bounded by a growing, jittered open period.
- Breaker state is observable next to the dependency's error rate.
Diagnostic Hook: chart breaker state and admitted traffic fraction against the dependency's error rate during incidents. A breaker that closes and re-opens repeatedly right after recovery is overloading a fragile dependency — switch to a ramp; a breaker that stays half-open at a low fraction long after the dependency is healthy has steps that are too long.
Pitfalls & edge cases¶
- Unlimited concurrent probes. Without an in-flight flag, every waiting request becomes a probe.
- Closing fully on one success. Simulated: 118 errors and 2 re-trips against a fragile recovery.
- Health checks that do not exercise the real path. The breaker closes onto a still-broken operation.
- Fixed short open periods. A long outage gets probed every few seconds by every process.
Frequently Asked Questions¶
What is the half-open state of a circuit breaker?
The state after the open period ends, in which the breaker lets limited traffic through to test whether the dependency has recovered before returning to normal.
How many requests should a half-open circuit breaker allow?
One at a time for dependencies that recover at full capacity; a ramp of increasing fractions for fragile ones. In simulation, a ramp caused 36 errors against 118 for a single probe that then closed fully.
Should circuit breaker probes use health checks or real requests?
Health checks avoid risking user requests but may miss failures on the real path. Use a health check to leave the open state, then ramp real traffic to confirm.
How long should a circuit breaker stay open?
Start short, a few seconds, and grow the open period with jitter while probes keep failing, resetting it after recovery.
Related¶
- Circuit Breakers & Bulkheads — up to the topic overview.
- Combining circuit breakers with retries — how retries interact with open and half-open states.
- Resilience, Cancellation & Error Handling — the section overview.