Skip to content

Testing Async Circuit Breakers

A circuit breaker is mostly untested code paths: it only does something when a dependency is failing, and its interesting moments happen at a timeout boundary or under concurrent load. That makes it a good candidate for tests that control time and concurrency precisely. Measured on Python 3.14 with pytest 9, against a small asyncio breaker with an injectable clock: five tests using a fake clock — threshold, reset of the failure count, the exact reset boundary, a single probe under 50 concurrent callers, and a randomized invariant over 200 seeded scenarios — ran in 0.17 s. One real-time test of the reset took 1.05 s on its own. To check the tests rather than trust them, four bugs were seeded into the breaker one at a time: an off-by-one threshold, a failure count not reset on success, a reset one second early, and unlimited probes when half-open. The fake-clock suite caught all four; the real-time test caught two. Two of the bugs were each caught by one test only. This guide builds those tests.

Prerequisites

1. Inject the clock

The breaker reads time in one place — when deciding whether an open circuit may try again. Make that clock a parameter:

class Breaker:
    def __init__(self, threshold=5, reset_after=30.0, clock=time.monotonic):
        self.threshold, self.reset_after, self.clock = threshold, reset_after, clock
        self.state, self.failures, self.opened_at = "closed", 0, 0.0
        self.probe_in_flight = False

    async def call(self, fn, *args):
        if self.state == "open":
            if self.clock() - self.opened_at < self.reset_after:
                raise Open()
            self.state = "half_open"
        if self.state == "half_open":
            if self.probe_in_flight:
                raise Open()
            self.probe_in_flight = True
        try:
            result = await fn(*args)
        except Exception:
            self._failure()
            raise
        else:
            self._success()
            return result
        finally:
            self.probe_in_flight = False

Tests pass a clock object whose __call__ returns a number they control. Unlike a fake sleep, the breaker never waits, so the test simply moves the clock forward between calls. A real-time test of a 1-second reset, by contrast, must actually sleep: it took 1.05 s for one assertion, and production resets are tens of seconds.

Verify: no test of the breaker calls asyncio.sleep to wait for a reset.

2. Test the transitions and the boundary

Each state transition gets a test, and the reset gets a test on both sides of its boundary:

class Clock:
    def __init__(self):
        self.t = 1000.0
    def __call__(self):
        return self.t

def test_half_open_exactly_at_reset():
    async def scenario():
        clock = Clock()
        b = Breaker(threshold=1, reset_after=30, clock=clock)
        with pytest.raises(ConnectionError):
            await b.call(failing)
        clock.t += 29.999
        with pytest.raises(Open):
            await b.call(succeeding)             # still open just before the boundary
        clock.t += 0.001
        assert await b.call(succeeding) == "ok"  # probe allowed at the boundary
        assert b.state == "closed"
    asyncio.run(scenario())

A test of the threshold asserts the state is still closed after threshold - 1 failures and open after threshold; a test of recovery asserts that a success resets the failure count. Measured against the seeded bugs: the off-by-one threshold failed the threshold test; the missing reset on success failed only the reset-count test; the early reset failed the boundary test.

Verify: each transition — closed to open, open to half-open, half-open to closed and back to open — has a test that would fail if the transition happened one step early or late.

Which tests caught which seeded bug A grid of 4 rows by 3 columns. Which tests caught which seeded bug seeded bug fake-clock tests that failed real-time test threshold off by one threshold, boundary, single probe failed count not reset on success reset-count only passed reset 1 s early boundary, random invariant failed unlimited probes when half-open single probe (50 concurrent) only passed The fake-clock suite caught 4 of 4; the real-time test 2 of 4.

3. Test half-open under concurrency

The half-open state is where breakers fail in production: many callers are waiting when the reset elapses. Test it with many concurrent callers and count how many reach the dependency:

def test_one_probe_when_half_open():
    async def scenario():
        clock = Clock()
        b = Breaker(threshold=1, reset_after=30, clock=clock)
        with pytest.raises(ConnectionError):
            await b.call(failing)
        clock.t += 30
        reached = 0
        async def counted():
            nonlocal reached
            reached += 1
            await asyncio.sleep(0.01)            # the probe takes time; others arrive meanwhile
            raise ConnectionError()
        await asyncio.gather(*(b.call(counted) for _ in range(50)), return_exceptions=True)
        assert reached == 1
        assert b.state == "open"
    asyncio.run(scenario())

The small real sleep inside the probe matters: without it, the probe completes before the other callers run, and a breaker with no probe limit passes. Measured: the seeded "unlimited probes" bug was caught by this test and by no other — the same weakness found in three real libraries in using circuit breaker libraries.

Verify: the concurrency test fails when the probe limit is removed from the breaker.

4. Check an invariant over random scenarios

Hand-written cases cover the transitions you thought of. A randomized test checks one property across many sequences: while the breaker is open and the reset has not elapsed, no call reaches the dependency:

def test_never_calls_dependency_while_open():
    async def scenario(seed):
        r = random.Random(seed)
        clock = Clock()
        b = Breaker(threshold=3, reset_after=10, clock=clock)
        for _ in range(300):
            clock.t += r.random() * 3
            healthy = r.random() < 0.5
            was_open, opened_at = b.state == "open", b.opened_at
            reached = []
            async def dependency():
                reached.append(1)
                if not healthy:
                    raise ConnectionError()
            with contextlib.suppress(ConnectionError, Open):
                await b.call(dependency)
            if was_open and clock.t - opened_at < 10:
                assert not reached, seed
    for seed in range(200):
        asyncio.run(scenario(seed))

Measured: 200 scenarios of 300 calls each ran in 0.04 s, and the test caught the seeded early-reset bug. Assertions carry the seed, so a failure is reproducible by number. For shrinking failures to a minimal sequence, the same property can be driven by Hypothesis, as in property-based testing of async code with Hypothesis.

Verify: the invariant test fails on at least one deliberately broken breaker.

Test time 2 horizontal bars comparing 5 fake-clock tests (incl. 200 random scenarios) with the others. Test time 5 fake-clock tests (incl. 200 random scenarios) 0.17 s 1 real-time reset test 1.05 s The faster suite also caught more bugs.

5. Seed bugs to test the tests

A breaker test suite that has never failed proves little. Seed bugs behind an environment variable — or with a mutation-testing tool — and confirm each is caught:

BUG = os.environ.get("BUG", "")

def _failure(self):
    ...
    if self.failures >= self.threshold + (1 if BUG == "off_by_one" else 0):
        self._open()
for bug in off_by_one no_reset_on_success early_reset many_probes; do
    BUG=$bug pytest -q tests/test_breaker.py
done

Measured: each of the four bugs failed at least one fake-clock test; two were caught by exactly one test, so removing that test would have let the bug through. The real-time test missed the two bugs that did not involve its single reset path. Keep the seeding hooks out of production code — a test-only subclass or a mutation tool such as mutmut does the same job without them. The same discipline applies to retry logic in testing rate limiters deterministically.

Verify: a table of seeded bugs against failing tests shows every bug caught.

Building a breaker test suite A flow of 5 stages. Building a breaker test suite Inject the clock no real sleeps Transitions both sides of each boundary Half-open 50 callers, slow probe Invariant random seeded scenarios Seed bugs every one must fail a test Measured suite: 0.17 s, 4 of 4 seeded bugs caught.

Verification

A circuit breaker is well tested when:

  • Time is injected, and no test sleeps for a reset.
  • Every transition is tested on both sides of its threshold or boundary.
  • Half-open is tested under concurrency, with a probe slow enough for others to arrive.
  • Seeded bugs each fail at least one test.

Diagnostic Hook: when a breaker passes its tests but floods a recovering dependency in production, check whether any test calls it concurrently while half-open. The unlimited-probe bug in this experiment was caught only by the test with 50 concurrent callers and a probe that took 10 ms.

Pitfalls & edge cases

  • Real-time reset tests. Measured: 1.05 s per test, and 2 of 4 bugs missed.
  • Concurrency tests with an instant probe. The probe finishes before other callers run.
  • Suites never seen to fail. Two seeded bugs were each caught by a single test.
  • Seeding hooks in production code. Use a subclass or a mutation tool instead.

Frequently Asked Questions

How do I test a circuit breaker without waiting for its timeout?

Pass the breaker a clock function and advance it in the test. Five fake-clock tests ran in 0.17 s; one real-time test of a 1 s reset took 1.05 s.

How do I test the half-open state?

Advance the clock past the reset, then call the breaker from 50 concurrent tasks with a probe that awaits briefly, and assert exactly one reached the dependency.

How do I know my circuit breaker tests are good enough?

Seed bugs and check each fails a test. The fake-clock suite caught 4 of 4 seeded bugs; a real-time test caught 2.

What property should a circuit breaker always satisfy?

While open and before the reset elapses, no call reaches the dependency. A random test checked it across 200 seeded scenarios in 0.04 s.