Skip to content

Probing Strategies for Half-Open Circuit Breakers

When a circuit breaker's open period ends, it moves to half-open and lets some traffic through to test whether the dependency has recovered. The common textbook version allows a single probe and closes completely on success. That works for a dependency that comes back at full strength, and it is exactly wrong for one that comes back fragile — a database still warming its caches, a service that just restarted with cold connection pools, a provider recovering from overload. Simulated with 100 requests per second against a dependency that, for ten seconds after an outage, could handle only 30 per second: closing on the first successful probe sent full traffic at once, caused 118 errors after recovery and 2 re-trips; ramping admitted traffic through 10%, 25%, 50% and 100%, two seconds per step, caused 36 errors and 1 re-trip — at the cost of reaching full traffic 17 s after recovery instead of 13 s. This guide implements both, and chooses between them.

Prerequisites

1. Understand what half-open is for

Half-open answers one question — "is the dependency usable again?" — while risking as little as possible:

class Breaker:
    def __init__(self, open_for: float) -> None:
        self.state = "closed"
        self.open_for = open_for
        self.reopen_at = 0.0

    def admit(self, now: float) -> bool:
        if self.state == "open":
            if now < self.reopen_at:
                return False                         # fail fast
            self.state = "half_open"                 # time to test
        if self.state == "half_open":
            return self.half_open_admit(now)         # the strategy lives here
        return True

The open period gives the dependency time to recover without your traffic; the half-open policy decides how much traffic tests it, and how quickly to return to normal. Too little probing keeps users on the fallback longer than necessary; too much turns your recovering traffic into the next outage. The right amount depends on how your dependencies recover, which is worth observing in past incidents.

Verify: the breaker exposes its state as a metric, and half-open periods are visible on dashboards alongside the dependency's error rate.

2. Use a single probe for dependencies that recover cleanly

Admit one request at a time while half-open; close on success, reopen on failure:

class SingleProbe(Breaker):
    def __init__(self, open_for: float) -> None:
        super().__init__(open_for)
        self.probe_in_flight = False

    def half_open_admit(self, now: float) -> bool:
        if self.probe_in_flight:
            return False                           # everyone else still fails fast
        self.probe_in_flight = True
        return True

    def record(self, now: float, ok: bool) -> None:
        if self.state == "half_open":
            self.probe_in_flight = False
            if ok:
                self.state = "closed"              # full traffic immediately
            else:
                self.state, self.reopen_at = "open", now + self.open_for

Allowing exactly one probe matters in async code: without the probe_in_flight flag, every request that arrives while the first probe is waiting also becomes a probe. Simulated against a fragile recovery, the single probe succeeded as soon as the dependency answered at all, the breaker closed, 100 requests per second hit a dependency that could take 30, errors pushed the window over threshold, and the breaker re-opened — twice — for 118 errors in total. Against a dependency that recovers at full capacity it is ideal: one request of risk, then immediate full service.

Verify: while half-open, exactly one request reaches the dependency at a time.

Half-open strategies against a fragile recovery A grid of 2 rows by 4 columns. Half-open strategies against a fragile recovery strategy errors after recovery re-trips full traffic after single probe, then close 118 2 13 s ramp 10 / 25 / 50 / 100 %, 2 s per step 36 1 17 s 100 requests/s offered; dependency capacity 30/s for 10 s after recovery.

3. Ramp traffic for dependencies that recover fragile

Admit a growing fraction of traffic in steps, checking the error rate at each step before moving on:

import random


class RampingProbe(Breaker):
    STEPS = (0.10, 0.25, 0.50, 1.00)

    def __init__(self, open_for: float, step_seconds: float = 2.0, max_error_rate: float = 0.5) -> None:
        super().__init__(open_for)
        self.step_seconds, self.max_error_rate = step_seconds, max_error_rate
        self.step = 0
        self.step_started = 0.0
        self.ok = self.total = 0

    def half_open_admit(self, now: float) -> bool:
        if self.total == 0 and self.ok == 0 and self.step_started == 0.0:
            self.step_started = now
        return random.random() < self.STEPS[self.step]

    def record(self, now: float, ok: bool) -> None:
        if self.state != "half_open":
            return
        self.total += 1
        self.ok += ok
        if self.total >= 10 and (self.total - self.ok) / self.total >= self.max_error_rate:
            self.state, self.reopen_at = "open", now + self.open_for      # back off again
            self._reset()
        elif now - self.step_started >= self.step_seconds:
            if self.step == len(self.STEPS) - 1:
                self.state = "closed"
                self._reset()
            else:
                self.step += 1
                self.step_started, self.ok, self.total = now, 0, 0

    def _reset(self) -> None:
        self.step, self.step_started, self.ok, self.total = 0, 0.0, 0, 0

Simulated: the ramp held load at 10 and 25 requests per second during the fragile period — inside the dependency's 30 — re-tripped once when it stepped to 50%, and caused 36 errors against 118. Full traffic resumed after 17 s instead of 13 s, because each step takes time. Rejected requests during the ramp still get the fallback path, so users see degraded service rather than errors.

Verify: during a staged recovery in a test, admitted traffic follows the steps and the dependency's error rate stays under the step threshold.

Admitted traffic after recovery, by strategy 2 lanes over time. Admitted traffic after recovery, by strategy single probe 100% open 100% open 100% ramp 10% 25% 50% open ramp again time after the outage ends (not to scale) → The ramp keeps load near what a recovering dependency can take.

4. Probe with real requests or a health check

Probes can be ordinary user requests, as above, or dedicated health checks sent by the breaker itself:

async def health_probe_loop(breaker: Breaker, client, url: str, interval: float = 1.0) -> None:
    while True:
        await asyncio.sleep(interval)
        if breaker.state != "open" or time.monotonic() < breaker.reopen_at:
            continue
        try:
            async with asyncio.timeout(2.0):
                response = await client.get(url)
            healthy = response.status_code == 200
        except (TimeoutError, httpx.HTTPError):
            healthy = False
        if healthy:
            breaker.state = "half_open"            # now let real traffic ramp in
        else:
            breaker.reopen_at = time.monotonic() + breaker.open_for

Health-check probes cost nothing for users — no real request is risked — but a health endpoint can report healthy while the real operation still fails (a database reachable but a table locked, a service up but its own dependency down). The combination works well: use a health check to leave the open state, then ramp real traffic to confirm. With several processes, each has its own breaker and sends its own probes; share state as in sharing circuit breaker state across processes, or accept N probes for N processes.

Verify: while the dependency is down, probe traffic is bounded to one request per interval per process.

5. Back off the open period on repeated failures

A fixed open period re-tests a long outage every few seconds forever. Grow the open period while probes keep failing, and reset it on recovery:

class BackoffOpen:
    def __init__(self, base: float = 2.0, cap: float = 60.0) -> None:
        self.base, self.cap, self.failures = base, cap, 0

    def next_open_for(self) -> float:
        self.failures += 1
        delay = min(self.cap, self.base * 2 ** (self.failures - 1))
        return delay * random.uniform(0.8, 1.2)            # jitter across processes

    def reset(self) -> None:
        self.failures = 0

Short initial open periods recover quickly from brief blips; growing ones stop a fleet of processes from probing a dead dependency every two seconds for an hour. Jitter spreads the probes from many processes so they do not arrive together. The same exponential-backoff reasoning applies as in exponential backoff with jitter in asyncio.

Verify: during a long outage in a test, the interval between probes grows to the cap; after recovery, the next trip starts again at the base.

Which half-open strategy fits this dependency? A decision on How does the dependency come back with 4 outcomes. Which half-open strategy fits this dependency? How does the dependency come back? at full strength single probe, close on success fastest fragile, cold caches/pools ramp 10-25-50-100% 36 vs 118 errors has a real health check health probe, then ramp no user risk outages can be long back off the open period jittered Probe gently enough not to cause the next outage.

Verification

Half-open behaviour is sound when:

  • Only one probe is in flight at a time for single-probe breakers.
  • Fragile dependencies are ramped, with an error check at each step.
  • Probing load during long outages is bounded by a growing, jittered open period.
  • Breaker state is observable next to the dependency's error rate.

Diagnostic Hook: chart breaker state and admitted traffic fraction against the dependency's error rate during incidents. A breaker that closes and re-opens repeatedly right after recovery is overloading a fragile dependency — switch to a ramp; a breaker that stays half-open at a low fraction long after the dependency is healthy has steps that are too long.

Pitfalls & edge cases

  • Unlimited concurrent probes. Without an in-flight flag, every waiting request becomes a probe.
  • Closing fully on one success. Simulated: 118 errors and 2 re-trips against a fragile recovery.
  • Health checks that do not exercise the real path. The breaker closes onto a still-broken operation.
  • Fixed short open periods. A long outage gets probed every few seconds by every process.

Frequently Asked Questions

What is the half-open state of a circuit breaker?

The state after the open period ends, in which the breaker lets limited traffic through to test whether the dependency has recovered before returning to normal.

How many requests should a half-open circuit breaker allow?

One at a time for dependencies that recover at full capacity; a ramp of increasing fractions for fragile ones. In simulation, a ramp caused 36 errors against 118 for a single probe that then closed fully.

Should circuit breaker probes use health checks or real requests?

Health checks avoid risking user requests but may miss failures on the real path. Use a health check to leave the open state, then ramp real traffic to confirm.

How long should a circuit breaker stay open?

Start short, a few seconds, and grow the open period with jitter while probes keep failing, resetting it after recovery.