Testing Distributed Locks Under Pauses¶
A distributed lock with a lease is correct only while its holder finishes within the lease. Garbage collection, a blocking call on the event loop, CPU throttling or a VM migration can stall a holder past that point, at which moment the lease expires, another process takes the lock, and two processes believe they own it. Ordinary tests never see this, because nothing in them pauses. Measured with a Redis lock (SET NX PX, 500 ms lease) protecting a read-modify-write of a Postgres counter, three asyncio processes running for 10 s: with no pauses, 434 successful writes and a final counter of 434. With a chaos script stopping a random process with SIGSTOP for 0.8 s six times, three runs lost 4, 3 and 3 updates, with 42–53 critical sections overlapping another. A blocking 0.8 s sleep injected into every twentieth critical section lost 32 of 235. With a 2 s lease the same 0.8 s pauses lost nothing, and with a fencing token checked by Postgres, every run lost nothing — the stale writers' updates were rejected instead. A deterministic pytest that pauses a holder in-process failed 6 times out of 6 for the unfenced lock and passed for the fenced one, in 0.32 s per test. This guide builds both kinds of test.
Prerequisites¶
- Redis and Postgres, here in local Docker containers; asyncpg and redis-py 5+.
- The lock under test, from implementing a Redis lock with fencing tokens.
- The topic overview, Distributed Locks & Coordination.
1. State the invariant the lock protects¶
A lock test needs something to check that is not "the lock was acquired". Choose the property the lock exists for, and make it measurable. For a counter protected by a lock, every successful write must be reflected exactly once:
RELEASE = "if redis.call('GET', KEYS[1]) == ARGV[1] then return redis.call('DEL', KEYS[1]) end return 0"
async def increment(r, pg, lease_ms):
token = uuid.uuid4().hex
while not await r.set("lock:counter", token, nx=True, px=lease_ms):
await asyncio.sleep(0.005)
try:
value = await pg.fetchval("SELECT value FROM counter WHERE id = 1")
await asyncio.sleep(0.02) # the work
await pg.execute("UPDATE counter SET value = $1 WHERE id = 1", value + 1)
return True
finally:
await r.eval(RELEASE, 1, "lock:counter", token)
Each process counts the writes it made; at the end, the sum of reported writes must equal the counter. A second, independent check helps diagnose violations: record each critical section's start and end using the Redis server clock (TIME), then count sections that overlap another. Measured without any pauses: 434 writes, counter 434, zero overlaps. That baseline matters — a test harness that reports violations without faults is measuring itself.
Verify: the invariant holds over a fault-free run, and the checker reports violations when you deliberately break the lock (for example by removing nx=True).
2. Pause whole processes with SIGSTOP¶
The most realistic pause is the whole process stopping — what a long GC pause, a container freeze or a hypervisor stall looks like from outside. SIGSTOP produces exactly that, and SIGCONT resumes the process as if nothing happened:
# start three workers, then pause a random one for 0.8 s every 0.7 s
pids=()
for i in 1 2 3; do python worker.py & pids+=($!); done
end=$((SECONDS + 9))
while [ $SECONDS -lt $end ]; do
sleep 0.7
victim=${pids[$((RANDOM % 3))]}
kill -STOP "$victim" && sleep 0.8 && kill -CONT "$victim"
done
wait "${pids[@]}"
Measured with a 500 ms lease over three 10 s runs: 352, 379 and 376 successful writes against final counters of 348, 376 and 373 — 4, 3 and 3 lost updates — with 53, 42 and 42 overlapping critical sections. Only pauses that land while a process holds the lock cause damage, which is why six pauses produced three or four losses, and why such a test must run long enough to hit the window. Repeated with a 2 s lease and the same 0.8 s pauses: 439 writes, counter 439, no overlaps — a pause shorter than the lease is harmless.
Verify: the chaos run reports lost updates for a lease shorter than the injected pause, and none for a lease comfortably longer.
3. Inject pauses inside the process¶
SIGSTOP tests are realistic but random. To make violations frequent, inject the pause where it hurts — inside the critical section — with a hook that blocks the event loop the way a stray synchronous call would:
INJECT = int(os.environ.get("INJECT", "0"))
async def increment(r, pg, lease_ms, n):
...
value = await pg.fetchval("SELECT value FROM counter WHERE id = 1")
await asyncio.sleep(0.02)
if INJECT and n % INJECT == INJECT - 1:
time.sleep(0.8) # blocks the whole loop: a GC pause stand-in
await pg.execute("UPDATE counter SET value = $1 WHERE id = 1", value + 1)
Measured with a pause in every twentieth critical section: 235 successful writes, counter 203, 32 lost updates, 113 overlapping sections. Each injected pause guaranteed that the lease expired while the lock was held, so nearly every one produced a violation — and a blocked loop also delays the holder's own heartbeat or renewal task, which is why renewing the lease from the same event loop, as in renewing lock leases with a heartbeat task, does not protect against this kind of pause.
Verify: with injection on, the unfenced lock loses updates in every run.
4. Write a deterministic test for CI¶
Chaos runs take seconds and depend on timing; CI needs a test that fails the same way every time. Make the pause explicit: let one holder take the lock and then wait on an asyncio.Event, let the lease expire, let a second holder in, and only then release the first:
LEASE_MS = 200
async def increment(r, pg, fenced, pause=None):
token, fence = await acquire(r) # SET NX PX, then INCR for a fence
value = await pg.fetchval("SELECT value FROM counter WHERE id = 1")
if pause:
await pause.wait() # held here until the test says so
sql = "UPDATE counter SET value = $1, fence = $2 WHERE id = 1"
if fenced:
sql += " AND fence < $2"
status = await pg.execute(sql, value + 1, fence)
await r.eval(RELEASE, 1, "lock:t", token)
return status.endswith(" 1")
@pytest.mark.parametrize("fenced", [False, True])
async def test_paused_holder_does_not_lose_updates(env, fenced):
r, pg = env
pause = asyncio.Event()
paused = asyncio.create_task(increment(r, pg, fenced, pause))
await asyncio.sleep(LEASE_MS / 1000 * 1.5) # lease expires during the stall
assert await increment(r, pg, fenced) # a second holder gets in
pause.set() # the first one wakes up
first_wrote = await paused
value = await pg.fetchval("SELECT value FROM counter WHERE id = 1")
assert value == 1 + first_wrote
Measured over six runs: the unfenced case failed every time with "2 writes reported, counter is 1", and the fenced case passed every time; each case took 0.32 s against real Redis and Postgres. Because the pause is an await, the test is deterministic: the order of events is fixed by the test, not by the scheduler. Run the async tests with pytest-asyncio as in Testing Async Code.
Verify: the test fails against the unfenced lock and passes against the fenced one, on every run.
5. Fix what the tests find, and keep them running¶
The tests show what a lease cannot do: it cannot stop a paused holder from acting when it wakes. Longer leases make the window rarer — the 2 s lease survived 0.8 s pauses — but no lease outlasts every possible pause. The fix that made every test pass is to make the protected resource reject stale holders: each acquisition takes a monotonically increasing fence, and the write succeeds only if no newer fence has written:
UPDATE counter SET value = $1, fence = $2
WHERE id = 1 AND fence < $2; -- 0 rows: a newer holder has written; ours is stale
Measured with fencing: in the SIGSTOP run, 417 writes and counter 417, with 1 stale write rejected; with injected pauses, 212 and 212, with 10 rejected. A rejected write is not an error to retry blindly — the holder's view of the data is stale, so it must re-read under a new acquisition. Keep the deterministic test in CI and run the chaos script before changing lease durations, lock libraries or the storage layer, since any of them can reopen the window.
Verify: both the deterministic test and a chaos run pass with the lease and storage configuration used in production.
Verification¶
A distributed lock is tested against pauses when:
- An invariant — not lock acquisition — is checked after every run.
- A chaos run stops whole processes for longer than the lease.
- A deterministic test pauses a holder past its lease and runs in CI.
- The fenced version passes both, and rejected stale writes are re-read, not replayed.
Diagnostic Hook: in production, count fence rejections at the storage layer and export the count. Each rejection is a pause that outlived a lease and would have been a lost update without fencing; a rising count means leases are too short for the pauses your processes actually experience.
Pitfalls & edge cases¶
- Testing only acquisition. The lock always "works"; the data does not.
- Renewing from the paused loop. A blocked loop blocks its heartbeat too.
- Short chaos runs. Only pauses inside the critical section count.
- Retrying fence-rejected writes as-is. The value they carry is stale.
Frequently Asked Questions¶
How do I test a distributed lock?
Check the invariant the lock protects, such as no lost updates, while pausing holders past their lease. A SIGSTOP chaos run lost 3 to 4 updates per 10 s with a 500 ms lease in testing, and a deterministic pytest caught the same bug in 0.32 s.
How do I simulate a GC pause in a Python test?
Stop the process with SIGSTOP and resume it with SIGCONT, or call time.sleep inside the critical section to block the event loop. For deterministic tests, hold the holder on an asyncio.Event while the lease expires.
Does a longer lease fix lost updates?
It makes them rarer: a 2 s lease survived 0.8 s pauses with no losses. No lease covers every possible pause, so use fencing where correctness matters.
What should happen when a fenced write is rejected?
The holder's data is stale. Re-acquire, re-read and recompute; do not retry the same write.
Related¶
- Distributed Locks & Coordination — up to the topic overview.
- Distributed semaphores with Redis — the same lease problem for N holders.
- Concurrent Execution & Worker Patterns — the section overview.