Renewing Lock Leases with a Heartbeat Task¶
A distributed lock's lease must be short so a crashed holder does not block everyone for long, and the work it protects is often longer than any sensible lease. The bridge is renewal: while the work runs, a separate task extends the lease at regular intervals. The renewal is half of the pattern; the other half is what happens when renewal fails — because then you may no longer hold the lock, and continuing is exactly the double-processing the lock exists to prevent. Tested against Redis 7.4 with a 1.5 s lease renewed every 0.5 s, a job held the lock for 3 s — twice the lease — without losing it; when another worker took the key over, the heartbeat's next renewal failed and the job was cancelled, in that run 5 ms after the takeover (the worst case is one renewal interval). This guide builds the heartbeat with the right structure, intervals and failure handling.
Prerequisites¶
- Python 3.11+,
redis.asyncioand a Redis server (the pattern applies to any lease). - The lock itself, from implementing a Redis lock with fencing tokens.
- Structured concurrency, from structured concurrency with asyncio.TaskGroup.
1. Renew only if you still own the lease¶
Extension must be conditional on ownership, atomically, or a holder whose lease expired would extend someone else's lock:
EXTEND = """
if redis.call('get', KEYS[1]) == ARGV[1] then
return redis.call('pexpire', KEYS[1], ARGV[2])
else
return 0
end
"""
async def extend(r, key: str, token: str, ttl_ms: int) -> bool:
return bool(await r.eval(EXTEND, 1, key, token, ttl_ms))
A 0 result is unambiguous: the key is gone or belongs to someone else. Treat a failed call — a Redis timeout or connection error — as uncertain, and decide in step 4 how long uncertainty may last.
Verify: extend after another client has overwritten the key; the call returns False and the other client's TTL is unchanged.
2. Run the heartbeat beside the work, in one TaskGroup¶
The work and the heartbeat must live and die together. A TaskGroup gives that: if the heartbeat discovers the lease is lost, it raises, the group cancels the work, and the exception reaches the caller:
import asyncio
class LeaseLost(Exception):
pass
async def run_with_lease(r, key: str, token: str, ttl_ms: int, work) -> None:
interval = ttl_ms / 1000 / 3
async def heartbeat() -> None:
while True:
await asyncio.sleep(interval)
if not await extend(r, key, token, ttl_ms):
raise LeaseLost(key)
async with asyncio.TaskGroup() as tg:
hb = tg.create_task(heartbeat())
await work() # the group body runs the work
hb.cancel() # work finished first: stop renewing
Running the work in the group's body and the heartbeat as a child means a LeaseLost from the heartbeat cancels the body — the work — at its next await. When the work finishes normally, cancelling the heartbeat ends the group cleanly. Measured: a 1.5 s lease stayed held through 3 s of work; after a takeover, the heartbeat raised on its next renewal and the work was cancelled.
Verify: delete the key mid-run in a test; the caller receives LeaseLost (inside an ExceptionGroup) within one renewal interval.
3. Choose the interval from loop lag, not from the lease alone¶
Renewal must happen before the lease expires even when the event loop is slow. The margin is the lease minus the interval, and it must exceed the worst realistic delay to the heartbeat task:
| Lease | Interval (lease/3) | Margin | Survives loop stalls up to |
|---|---|---|---|
| 1.5 s | 0.5 s | 1.0 s | ~1 s |
| 10 s | 3.3 s | 6.7 s | ~6 s |
| 30 s | 10 s | 20 s | ~20 s |
If your service has occasional 2-second stalls — a big synchronous parse, a GC pause on a huge heap — a 1.5 s lease will be lost during them no matter how often you renew. Measure stalls with a lag monitor, as in measuring event loop lag in production, and size the lease from the worst observed stall, not the average. The cost of a longer lease is slower recovery when the holder genuinely dies.
Verify: run the heartbeat under realistic load and record the maximum gap between successful renewals; it must stay well below the lease.
4. Decide what an unreachable lock store means¶
A renewal can fail because the lease is lost (False) or because Redis did not answer (an exception). The second is ambiguous: the lease may still be valid. The safe rule is to stop working before the lease could have expired:
import time
async def heartbeat_with_deadline(r, key, token, ttl_ms) -> None:
interval = ttl_ms / 1000 / 3
valid_until = time.monotonic() + ttl_ms / 1000
while True:
await asyncio.sleep(interval)
try:
if not await extend(r, key, token, ttl_ms):
raise LeaseLost(key)
valid_until = time.monotonic() + ttl_ms / 1000
except (ConnectionError, TimeoutError, OSError):
if time.monotonic() > valid_until - interval: # could expire before we know
raise LeaseLost(f"{key}: lock store unreachable")
valid_until tracks the last moment the lease was confirmed; transient Redis errors are tolerated until there is no longer a safe margin. This keeps short Redis blips from killing long jobs while guaranteeing that a long outage stops the work before another holder could exist. Fencing tokens, from the lock guide, remain the backstop for the cases this logic cannot see — a process paused so long that even the heartbeat could not run.
Verify: block Redis for less than the margin — work continues; block it for longer — work stops before the lease would have expired.
5. Make the work stoppable at awaits¶
Cancelling the work only helps if it reaches an await soon. Long synchronous sections inside locked work cannot be interrupted, and a CPU loop with no awaits will keep running after the lease is lost:
async def rebuild_index(fence: int) -> None:
async for batch in source_batches():
rows = await asyncio.to_thread(transform, batch) # CPU off the loop
await write_rows(rows, fence=fence) # an await between batches
Batch the work so there is an await every few hundred milliseconds, move CPU-heavy steps off the loop, and attach the fence to every write so a write that slips through after cancellation is rejected. Cooperative cancellation in CPU loops is covered in implementing cooperative cancellation in CPU loops.
Verify: after a forced lease loss, no write with the old fence is accepted by the resource.
Verification¶
The heartbeat is correct when:
- Renewal is conditional on ownership and atomic.
- Work and heartbeat share a TaskGroup, so a lost lease cancels the work.
- The renewal margin exceeds the worst measured loop stall.
- Unreachable lock stores stop work before the lease could expire, and fences catch the rest.
Diagnostic Hook: export the gap between successful renewals as a histogram and alert when its maximum exceeds half the lease. That gap is the real safety margin in production; it shrinks long before leases are actually lost, so it warns about rising loop lag or a degrading lock store while there is still time to act.
Pitfalls & edge cases¶
- Heartbeat in a separate, unconnected task. It keeps renewing after the work has crashed, holding the lock for nothing.
- Renewal without an ownership check. Extends another holder's lock.
- Ignoring renewal failures. The work continues as a stale holder.
- Lease shorter than loop stalls. Locks are lost during every stall.
Frequently Asked Questions¶
How do I keep a Redis lock while a long task runs?
Run a heartbeat task beside the work that extends the lease every third of its TTL with a compare-and-extend script, both inside one TaskGroup, and stop the work if an extension fails. In testing, a 1.5 s lease was held for 3 s this way.
How often should a lock lease be renewed?
Roughly every third of the lease, and the remaining margin must exceed the worst event loop stall you observe, otherwise the lease expires during stalls regardless of the interval.
What should happen when lease renewal fails?
If the store says the lock is not yours, stop immediately. If the store is unreachable, continue only while the last confirmed renewal still guarantees the lease is valid, then stop.
Is renewal enough to prevent two lock holders?
No. A process paused longer than the lease cannot renew or notice. Fence every write with a token so the resource rejects a stale holder that resumes.
Related¶
- Distributed Locks & Coordination — up to the topic overview.
- Leader election for asyncio workers — renewal held indefinitely.
- Concurrent Execution & Worker Patterns — the section overview.