Property-Based Testing of Async Code with Hypothesis¶
Example-based tests check the cases you thought of; property-based tests check a rule against hundreds of generated inputs and, when one fails, shrink it to the smallest input that still fails. For async code the inputs worth generating are often sequences — of keys, of concurrent calls, of operations — because that is where interleaving bugs hide. Tested with Hypothesis 6.168 and pytest-asyncio 1.4: @given worked directly on async def tests. A "single-flight" helper that shares one in-flight fetch between concurrent callers had a bug in its cleanup: it removed the in-flight entry only when other keys were also in flight. A property — each round of concurrent calls fetches each distinct key exactly once — caught it, and Hypothesis shrank the failure to rounds=[['a'], ['a']]: one key, requested once, then once more, and the second round returned the first round's cached result without fetching. The whole run took 0.34 s; the fixed version passed 200 generated examples. This guide writes properties for async components and keeps them fast and reproducible.
Prerequisites¶
- Python 3.11+,
pip install hypothesis pytest pytest-asyncio(tested with 6.168, 9.1, 1.4). - pytest-asyncio setup, from testing asyncio code with pytest-asyncio.
- Race reproduction, from reproducing race conditions deterministically.
1. State the property, not the examples¶
A property is something that must hold for every input. For a single-flight helper — concurrent get(key) calls share one fetch — a good one is: in each round of concurrent calls, the number of fetches equals the number of distinct keys:
from hypothesis import given, settings, strategies as st
rounds = st.lists(st.lists(st.sampled_from("abc"), min_size=1, max_size=6), min_size=1, max_size=5)
@settings(max_examples=200, deadline=None)
@given(rounds=rounds)
async def test_one_fetch_per_key_per_round(rounds):
fetches = []
async def fetch(key):
await asyncio.sleep(0)
fetches.append(key)
return key
sf = SingleFlight(fetch)
for keys in rounds:
before = len(fetches)
results = await asyncio.gather(*(sf.get(k) for k in keys))
assert results == keys
assert len(fetches) - before == len(set(keys))
The strategy generates rounds of concurrent requests over a small key space, so duplicates within a round (which must share a fetch) and repeats across rounds (which must fetch again) both occur often. deadline=None avoids spurious failures from Hypothesis's per-example time limit on loaded CI machines. Tested: @given on an async def test ran under pytest-asyncio 1.4 without extra setup.
Verify: the property fails if the helper either fetches twice for one key in a round, or never fetches again in a later round.
2. Let shrinking explain the bug¶
When a property fails, Hypothesis searches for a smaller failing input and reports the smallest it finds:
E AssertionError: round ['a']: 0 fetches for 1 distinct keys
E Falsifying example: test_one_fetch_per_key_per_round(
E rounds=[['a'], ['a']],
E )
That was the tested output. The minimal case reads like a bug report: request key a, then request it again — and the second request was served from the first round's completed future, without a fetch. The cause was the cleanup condition (del inflight[key] only when other keys were in flight); with one key in flight, the completed future stayed in the map forever and every later call returned the stale result. A hand-written test with several keys at once would have passed. The fix is unconditional cleanup in finally, after which the property held for all 200 examples.
Verify: after fixing, rerun with --hypothesis-seed=<seed> from the failure report to confirm the exact failing case now passes.
3. Generate schedules, not just data¶
Many async bugs depend on order and timing rather than values. Make the interleaving part of the generated input, with delays drawn by Hypothesis so failures shrink and replay:
@settings(max_examples=300, deadline=None)
@given(delays=st.lists(st.sampled_from([0, 0, 1, 2]), min_size=2, max_size=6))
async def test_withdrawals_never_overdraw(delays):
store = Store(balance=100)
async def withdraw_after(ticks: int) -> bool:
for _ in range(ticks):
await asyncio.sleep(0) # yield to the loop a chosen number of times
return await account.withdraw(store, 30)
results = await asyncio.gather(*(withdraw_after(d) for d in delays))
assert store.balance >= 0
assert store.balance == 100 - 30 * sum(results)
Yielding with asyncio.sleep(0) a generated number of times moves each task's start relative to the others without real time, so tests stay fast and deterministic per example. Shrinking then reduces the schedule to the fewest tasks and yields that still break the invariant — a ready-made deterministic reproduction. Keep the yield counts small; most interleaving bugs need only a step or two of offset.
Verify: introducing an await between check and write in withdraw makes the property fail with a two-task schedule.
4. Model stateful components with a state machine¶
For components with many operations — a cache, a connection pool, a rate limiter — a rule-based state machine generates whole sequences of operations and compares the component with a simple model:
from hypothesis.stateful import RuleBasedStateMachine, invariant, rule
class CacheMachine(RuleBasedStateMachine):
def __init__(self) -> None:
super().__init__()
self.loop = asyncio.new_event_loop()
self.cache = AsyncTTLCache(maxsize=3)
self.model: dict[str, int] = {}
@rule(key=st.sampled_from("abcd"), value=st.integers())
def put(self, key, value):
self.loop.run_until_complete(self.cache.put(key, value))
self.model[key] = value
if len(self.model) > 3:
self.model.pop(next(iter(self.model))) # model of the eviction policy
@invariant()
def matches_model(self):
for key, value in self.model.items():
assert self.loop.run_until_complete(self.cache.get(key)) == value
def teardown(self) -> None:
self.loop.close()
TestCache = CacheMachine.TestCase
State machines run synchronously, so each rule drives the async component with its own loop (run_until_complete). The model is deliberately simple — a dict and an eviction rule — and the invariant compares the two after every step; Hypothesis shrinks a failing sequence to the shortest run of operations that breaks it.
Verify: a deliberate off-by-one in the cache's eviction makes the machine fail with a short sequence of put calls.
5. Keep property tests fast and reproducible¶
Properties run many examples, so each must be quick, and failures must replay exactly:
from hypothesis import settings
settings.register_profile("ci", max_examples=500, deadline=None, derandomize=False,
print_blob=True)
settings.register_profile("dev", max_examples=100, deadline=None)
settings.load_profile(os.environ.get("HYPOTHESIS_PROFILE", "dev"))
Avoid real sleeps and real I/O inside properties — use fakes and asyncio.sleep(0) yields — so examples take microseconds; the tested examples ran in under a millisecond each. Hypothesis stores failing examples in its database (.hypothesis/) and replays them first on the next run; print_blob=True prints a reproduction blob for CI failures that can be pasted into @reproduce_failure. Turn shrunk counterexamples into ordinary example tests once fixed, so the regression stays covered even if the property's strategy changes.
Verify: a CI failure can be reproduced locally from its printed blob or seed in one run.
Verification¶
Property-based tests are effective when:
- Properties state rules, not single expected outputs.
- Strategies generate concurrency patterns, including repeats and duplicates.
- Schedules and operation sequences are generated where order matters.
- Examples are fast and failures reproducible, and fixed counterexamples become example tests.
Diagnostic Hook: run with --hypothesis-show-statistics to see how many examples ran, how long they took and how often strategies hit interesting cases. Properties that spend most of their budget on trivial inputs need strategies weighted toward collisions — small key spaces, short lists — which is where async bugs live.
Pitfalls & edge cases¶
- Large key spaces. Collisions become rare and concurrency bugs go unexercised.
- Real sleeps in properties. Hundreds of examples become minutes.
- Assertions that only check results. The tested bug returned correct-looking values; counting fetches caught it.
- Leaving the default per-example deadline in CI. Loaded machines produce false failures.
Frequently Asked Questions¶
Does Hypothesis work with async test functions?
Yes with pytest-asyncio: @given on an async def test ran directly in testing with Hypothesis 6.168 and pytest-asyncio 1.4.
How do I use property-based testing to find concurrency bugs?
Generate the concurrency itself: lists of concurrent requests over a small key space, or per-task counts of asyncio.sleep(0) yields, and assert invariants that must hold for every interleaving.
What is a good property for an async cache or single-flight helper?
That each round of concurrent requests triggers exactly one fetch per distinct key. In testing it exposed a stale-result bug and shrank it to [['a'], ['a']].
How do I test a stateful async component with Hypothesis?
Use a RuleBasedStateMachine whose rules drive the component through its own event loop with run_until_complete and an invariant that compares it with a simple model.
Related¶
- Testing Async Code — up to the topic overview.
- Testing with unittest IsolatedAsyncioTestCase — the stdlib alternative for example tests.
- Resilience, Cancellation & Error Handling — the section overview.