Reproducing Race Conditions Deterministically¶
asyncio races happen at await points: one task reads some state, awaits, and acts on what it read while another task changed it in between. Whether a given run interleaves badly depends on timing, so the bug is intermittent — and intermittent bugs survive test suites. Measured with a classic check-then-act withdrawal (read the balance, await, write the new balance), two concurrent withdrawals of 80 from a balance of 100 overdrew the account in 537 of 1,000 runs when the read always awaited briefly; when the read awaited only 1% of the time — a cache miss — it overdrew in 107 of 10,000 runs (1.07%), the kind of failure that appears once a week in CI and never on a laptop. Forcing the interleaving with an asyncio.Event that held the first reader until the second had also read made the failure happen in 100 of 100 runs; adding a lock around the read-modify-write made it 0. This guide turns races into deterministic tests and then fixes them.
Prerequisites¶
- Python 3.11+,
pip install pytest pytest-asyncio. - Synchronization primitives, from choosing asyncio Lock vs Semaphore vs Event.
- Mocking async dependencies, from mocking async dependencies with AsyncMock.
1. Find the await between check and act¶
A race in single-threaded asyncio needs an await between reading state and acting on it. Everything between two awaits runs without interruption:
async def withdraw(store, amount: int) -> bool:
balance = await store.read() # check ... (another task can run here)
if balance >= amount:
await store.write(balance - amount) # ... act on a value that may be stale
return True
return False
# Two concurrent withdraw(store, 80) from balance 100:
# A reads 100 | B reads 100 | A writes 20 | B writes 20 -> both succeed: 160 withdrawn
Measured: with an await in every read, 537 of 1,000 runs overdrew; with an await in 1% of reads, 107 of 10,000. The shape — read, await, decide, write — is the thing to look for in code review: balance and inventory updates, "create if not exists", rate-limit counters, cache fills. The await may be hidden inside a helper, which is why the 1% case is the realistic one: the cache usually answers synchronously, and the race only opens on a miss.
Verify: list the awaits between each read and the write that depends on it; every one is a potential interleaving point.
2. Force the interleaving with a gate¶
Make the bad schedule happen every time by controlling the await point. Replace the dependency's slow step with a gate that releases only when both tasks have reached it:
async def test_concurrent_withdrawals_cannot_overdraw():
store = Store(balance=100)
arrived = 0
both_read = asyncio.Event()
real_read = store.read
async def gated_read():
nonlocal arrived
value = await real_read() # take the snapshot...
arrived += 1
if arrived == 2:
both_read.set()
await asyncio.wait_for(both_read.wait(), 0.5) # ...and hold it until the other task has one too
return value
store.read = gated_read
results = await asyncio.gather(withdraw(store, 80), withdraw(store, 80))
assert sum(results) == 1, "only one withdrawal may succeed"
Measured: with the gate, the overdraft happened in 100 of 100 runs — the test fails deterministically on the buggy code. The wait_for timeout matters: once the code is fixed with a lock, the second task cannot reach the gate while the first holds the lock, and the timeout lets the test proceed instead of deadlocking (the fixed version then shows one success, as intended). Gates like this can be placed by patching a dependency, by an AsyncMock side effect, or by a test hook in the code.
Verify: the gated test fails against the unfixed code on every run.
3. Fix the race, then keep the test¶
Once the test fails reliably, fix the code so the check and the act cannot be separated by another task:
class Account:
def __init__(self, store) -> None:
self.store = store
self.lock = asyncio.Lock()
async def withdraw(self, amount: int) -> bool:
async with self.lock: # no other withdraw can interleave
balance = await self.store.read()
if balance < amount:
return False
await self.store.write(balance - amount)
return True
Measured: with the lock, 0 of 1,000 random runs and 0 of 20 forced runs overdrew. A lock serializes the operation within one process; with several processes or services touching the same data, the fix moves to the data store — a conditional update (UPDATE ... SET balance = balance - $1 WHERE balance >= $1), a transaction with the right isolation, or a distributed lock — and the gated test still applies to the code that issues it. Keep the deterministic test in the suite: it is the regression test for the race.
Verify: the gated test passes on the fixed code, and fails again if the lock is removed.
4. Explore schedules when you cannot place a gate¶
When the racing code is spread across many awaits and you do not know which pair interleaves badly, randomize the schedule instead and run many iterations:
import random
class JitteredStore(Store):
async def read(self):
value = await super().read()
await asyncio.sleep(random.uniform(0, 0.001)) # widen every await point
return value
@pytest.mark.parametrize("seed", range(200))
async def test_withdraw_race_fuzzed(seed):
random.seed(seed)
store = JitteredStore(balance=100)
results = await asyncio.gather(*(withdraw(store, 80) for _ in range(3)))
assert sum(results) <= 1
Injecting small random delays at await points turns rare interleavings into common ones; parametrizing by seed makes every failure reproducible by its seed. This is a search tool — once it finds a failing seed, read the interleaving it produced and write a gated test for it. Property-based testing generalizes the idea, as in property-based testing of async code with Hypothesis.
Verify: a failing seed reproduces the failure every time it is rerun.
5. Make CI catch the rest¶
Some races only appear under load or on slow machines. Run concurrency-sensitive tests in ways that shake schedules:
pytest -p no:randomly tests/concurrency --count=50 # pytest-repeat: many iterations
PYTHONASYNCIODEBUG=1 pytest tests/concurrency # slow-callback warnings, stricter checks
pytest tests/concurrency -n 8 # parallel workers add CPU contention
Repeating tests turns low-probability failures into likely ones within a CI run; asyncio's debug mode adds checks and slows callbacks slightly, which shifts interleavings; running under CPU contention changes timing further. A test that fails once in 50 repetitions is a race until proven otherwise — capture the seed or the interleaving, then make it deterministic with a gate.
Verify: concurrency tests run repeatedly in at least one CI job, and any failure is triaged as a potential race.
Verification¶
Races are under control when:
- Every read-await-write sequence on shared state is reviewed for interleaving.
- Known races have gated tests that fail deterministically without the fix.
- Unknown races are searched with seeded random delays.
- CI repeats concurrency tests, and flakes are treated as races.
Diagnostic Hook: in production, assert invariants where they are cheap — balances never negative, counters never above limits — and log violations with the involved request ids. Each violation is a race report with its interleaving, which a gated test can then reproduce.
Pitfalls & edge cases¶
- Calling a 1% failure "flaky". Measured: a real race at 1.07%.
- Gates without timeouts. The fixed code deadlocks the test.
- Locks across processes. An
asyncio.Lockonly covers one event loop. - Deleting the gated test after the fix. It is the regression test.
Frequently Asked Questions¶
How do race conditions happen in single-threaded asyncio?
At await points: a task reads state, awaits, and acts on it while another task changed the state in between. In testing, two concurrent withdrawals overdrew an account in 537 of 1,000 runs.
How do I reproduce an asyncio race condition reliably?
Gate the await point: replace the awaited dependency with one that waits on an asyncio.Event until the other task has also reached it. That made the race fail in 100 of 100 runs.
How do I fix a check-then-act race in asyncio?
Hold an asyncio.Lock around the read, decision and write within one process, or use an atomic conditional update in the database when several processes share the data.
How do I find races when I don't know where they are?
Inject small random delays at await points, run many seeded iterations, and turn any failing seed into a deterministic gated test.
Related¶
- Testing Async Code — up to the topic overview.
- Detecting leaked tasks in tests — another bug class that passing tests hide.
- Resilience, Cancellation & Error Handling — the section overview.