Skip to content

Testing Rate Limiters Deterministically

A rate limiter is code whose whole job is about time, which makes it awkward to test: a test that measures real elapsed time is slow, and it fails whenever the machine is busy. The fix is to give the limiter its clock and its sleep function as parameters, then drive it with virtual time in tests. Measured on Python 3.14 with pytest 9 and a token bucket allowing 50 acquisitions per second: a real-time test asserting that 51 acquisitions take 1.00 s ± 1% passed 10 of 10 runs on an idle machine, taking about 1.1 s each, and failed 10 of 10 runs with eight busy processes on the same core, where the same acquisitions took 2.38–3.61 s. With an injected fake clock, the same check asserted exact timestamps — every gap 0.02 s to within 1e-9 — and a second test ran 300 randomized multi-client scenarios against a rate invariant; together they took 0.29 s. That randomized test found a real bug that no real-time test had: in 198 of 300 scenarios the limiter looped forever, because a remaining token deficit of 2.2e-16 produced a sleep too small to move the clock. A deliberately broken limiter that forgot to cap stored tokens violated the invariant in 118 of 300 scenarios. This guide builds the test setup and shows both catches.

Prerequisites

1. See why wall-clock tests fail

The obvious test runs the limiter and times it:

def test_rate_real_time():
    async def run():
        b = TokenBucket(rate=50, burst=1)
        t = time.monotonic()
        for _ in range(51):
            await b.acquire()
        return time.monotonic() - t
    elapsed = asyncio.run(run())
    assert 0.99 <= elapsed <= 1.01, elapsed     # 50 refills at 20 ms = 1.00 s

The first acquisition uses the stored token, and the next 50 each wait 20 ms, so the expected time is exactly one second. Measured on an idle machine: 10 of 10 runs passed, each pytest run taking about 1.1 s. With eight CPU-bound processes pinned to the same core — a stand-in for a busy CI runner — 10 of 10 runs failed, the loop taking 2.38 to 3.61 s because each 20 ms sleep woke up late. Widening the tolerance does not fix this: a test loose enough to survive a busy runner is too loose to catch a limiter that is 50% too permissive, and every such test still costs real seconds.

Verify: run the timing test under load, for example with taskset and a few busy loops on one core, and watch it fail without any code change.

The same limiter check, three ways A grid of 3 rows by 3 columns. The same limiter check, three ways test idle machine 8 busy processes on the core real time, 51 acquisitions, +-1% 10/10 pass, ~1.1 s 10/10 fail, 2.38-3.61 s fake clock, exact stamps pass, exact to 1e-9 pass, unaffected fake clock, 300 random scenarios pass pass Both fake-clock tests together ran in 0.29 s. Python 3.14, pytest 9.

2. Inject the clock and the sleep

The limiter reads the time and sleeps in exactly two places, so make both parameters with production defaults:

class TokenBucket:
    def __init__(self, rate: float, burst: int,
                 clock=time.monotonic, sleep=asyncio.sleep):
        self.rate, self.capacity = rate, burst
        self.tokens = float(burst)
        self.clock, self.sleep = clock, sleep
        self.updated = clock()
        self._lock = asyncio.Lock()

    def _refill(self):
        now = self.clock()
        self.tokens = min(self.capacity, self.tokens + (now - self.updated) * self.rate)
        self.updated = now

    async def acquire(self):
        async with self._lock:
            while True:
                self._refill()
                if self.tokens >= 1:
                    self.tokens -= 1
                    return
                await self.sleep((1 - self.tokens) / self.rate)

Production code constructs TokenBucket(rate=50, burst=1) and never sees the extra parameters. Tests pass a fake clock whose sleep advances virtual time instantly:

class FakeClock:
    def __init__(self):
        self.now = 0.0

    def time(self):
        return self.now

    async def sleep(self, seconds):
        self.now += max(0.0, seconds)
        await asyncio.sleep(0)          # still yield, so other tasks interleave

The await asyncio.sleep(0) matters: without it, a fake sleep never gives other tasks a turn, and tests with several concurrent clients would run each client to completion in sequence, hiding interleavings that happen in production. Injection is narrower than patching time.monotonic globally — patching affects every library in the process, including asyncio's own scheduler, which also reads the monotonic clock.

Verify: the production constructor call is unchanged, and no test patches time or asyncio.sleep globally.

3. Assert exact timestamps

With virtual time, the test can record when each acquisition happened and assert the schedule exactly:

def test_rate_fake_clock():
    async def run():
        clock = FakeClock()
        b = TokenBucket(rate=50, burst=1, clock=clock.time, sleep=clock.sleep)
        stamps = []
        for _ in range(51):
            await b.acquire()
            stamps.append(clock.time())
        return stamps
    stamps = asyncio.run(run())
    assert stamps[0] == 0.0 and abs(stamps[-1] - 1.0) < 1e-9
    assert all(abs((b - a) - 0.02) < 1e-9 for a, b in zip(stamps, stamps[1:]))

This asserts something the real-time test could not: not just that the total was about one second, but that every gap was 20 ms — so a limiter that admitted a burst of 10 and then stalled would fail it, even though its total time would be right. The tolerance is 1e-9, set by floating-point rounding rather than by scheduler jitter. It passed under the same eight-process contention that failed every real-time run.

Verify: the test asserts every interval, not just the total, and its result does not change when the machine is loaded.

One acquisition under a fake clock A sequence of 7 messages between 4 participants. One acquisition under a fake clock test TokenBucket FakeClock event loop await acquire() time() -> tokens 0.0 await sleep(0.02) now += 0.02, no real wait asyncio.sleep(0): let others run time() -> tokens 1.0 return; test records now Virtual time advances only when the code under test sleeps.

4. Check an invariant across random scenarios

Exact-schedule tests cover the cases you thought of. A property test covers the rest: generate many random workloads and check the one thing every token bucket must guarantee — any time window of length w admits at most burst + rate × w acquisitions.

def test_never_exceeds_burst_plus_rate_random():
    async def scenario(seed):
        r = random.Random(seed)
        clock = FakeClock()
        rate, burst = r.choice([5, 10, 50]), r.choice([1, 3, 10])
        b = TokenBucket(rate=rate, burst=burst, clock=clock.time, sleep=clock.sleep)
        stamps = []

        async def client():
            for _ in range(r.randint(1, 30)):
                await clock.sleep(r.random() * 0.2)      # arbitrary arrival pattern
                await b.acquire()
                stamps.append(clock.time())

        await asyncio.gather(*(client() for _ in range(r.randint(1, 8))))
        stamps.sort()
        for i, t0 in enumerate(stamps):
            for w in (0.1, 0.5, 1.0):
                n = sum(1 for t in stamps[i:] if t <= t0 + w + 1e-9)
                assert n <= burst + rate * w + 1e-6, (seed, rate, burst, w, n)

    for seed in range(300):
        asyncio.run(scenario(seed))

Each scenario picks a rate, a burst size, one to eight concurrent clients and random gaps between their requests, all from a seed, so any failure is reproducible by its seed number. To check that the test can fail, it was run against a limiter with one line removed — the min(self.capacity, ...) cap, so idle time accumulates unlimited tokens. That broken version violated the invariant in 118 of 300 scenarios, in 0.13 s. The correct version passed all 300. For shrinking failures to a minimal case, the same scenario can be driven by Hypothesis, as in property-based testing of async code with Hypothesis.

Verify: the property test fails against a deliberately broken limiter, and every assertion message includes the seed.

5. Fix what the random scenarios found

The first run of the property test against the limiter from step 2 did not fail an assertion — it hung. Instrumenting the loop showed scenario seed 1 stuck with tokens = 0.9999999999999998, sleeping for 4.44e-17 seconds at virtual time 1.713. Adding 4.44e-17 to 1.713 does not change a 64-bit float, so the clock never advanced, the refill added nothing, and the loop spun forever. Across all 300 seeds, 198 got stuck this way. Real time hides the bug because a real sleep always overshoots, however small the request; on a fast enough clock, or a virtual one, the deficit never closes.

The fix is to accept a token count within rounding error of one:

    async def acquire(self):
        async with self._lock:
            while True:
                self._refill()
                if self.tokens >= 1 - 1e-9:          # tolerate float rounding
                    self.tokens = max(0.0, self.tokens - 1)
                    return
                await self.sleep((1 - self.tokens) / self.rate)

With that change, all 300 scenarios passed and the whole file, including the exact-timestamp test, ran in 0.29 s. A minimum sleep, such as max(deficit / rate, 1e-6), fixes it too. Keep the seed that exposed the bug as a named regression test, so the case stays covered if the scenario generator changes.

Verify: the property test runs under a timeout, so a future hang fails the run instead of stalling CI, and the seed that found this bug has its own test.

300 random scenarios against three limiters 3 horizontal bars comparing original, exact compare to 1 with the others. 300 random scenarios against three limiters original, exact compare to 1 198 hung missing capacity cap 118 violations fixed, 1e-9 tolerance 0 Scenarios: rate 5/10/50 per s, burst 1/3/10, 1-8 clients, random arrivals. The hang appeared only under virtual time; real sleeps always overshoot.

Verification

The rate limiter is tested deterministically when:

  • The clock and sleep are injected, with production defaults, and tests never patch them globally.
  • Schedules are asserted exactly, interval by interval, with a tolerance set by float rounding.
  • A randomized invariant test covers many rates, bursts and client counts, reports seeds, and has been seen to fail against a broken limiter.
  • The suite runs under a timeout, so a livelock fails instead of hanging.

Diagnostic Hook: when a fake-clock test hangs, print the sleep argument and the clock value at each loop iteration. A sleep smaller than the clock's float spacing — 4.44e-17 against a clock near 1.7 here — means the loop is waiting for time that the arithmetic can never deliver.

Pitfalls & edge cases

  • Tolerance-based timing tests. Measured: 10 of 10 failed under CPU contention.
  • A fake sleep that never yields. Concurrent clients then run one at a time.
  • Comparing float token counts exactly. Measured: 198 of 300 scenarios hung.
  • A property test that has never failed. Break the limiter on purpose once; here it caught 118 of 300.

Frequently Asked Questions

How do I test a rate limiter without sleeping?

Make the limiter take its clock and sleep function as parameters, and pass a fake clock whose sleep advances virtual time instantly and then yields with asyncio.sleep(0).

Why does my rate limiter test fail on CI?

Real sleeps wake late on a busy machine. A test expecting 1.00 s plus or minus 1% failed 10 of 10 runs with eight busy processes on the core, measuring 2.38 to 3.61 s.

What property should a token bucket test check?

Any window of length w admits at most burst plus rate times w acquisitions. Check it over many random seeded scenarios with varying rates, bursts and client counts.

Why does a token bucket loop forever under a fake clock?

Float rounding can leave a deficit like 2.2e-16 token, whose sleep is too small to change the clock. Accept tokens within 1e-9 of one, or enforce a minimum sleep.