Skip to content

Implementing a Redis Lock with Fencing Tokens

The textbook Redis lock — SET key value NX PX ttl to acquire, DEL key to release — has two failure modes that both show up under real asyncio load. The first is releasing someone else's lock: a holder whose lease expired calls DEL and removes the lock a different worker now holds. The second is worse: a holder that paused past its lease (a long GC cycle, a blocked event loop, a slow network call) wakes up and keeps writing after another worker has taken over. The first is fixed by releasing only if the stored value is yours; the second needs a fencing token. Tested against Redis 7.4, a worker whose 300 ms lease expired during a 400 ms pause lost the lock to a second worker; its later write carried fencing token 1 while the current holder's carried 2, and the storage layer rejected it; its release attempt left the new holder's lock intact. This guide builds each piece.

Prerequisites

1. Acquire with SET NX PX and a unique value

import uuid
import redis.asyncio as aioredis


class RedisLock:
    def __init__(self, r: aioredis.Redis, name: str, ttl_ms: int) -> None:
        self.r = r
        self.key = f"lock:{name}"
        self.ttl = ttl_ms
        self.token: str | None = None
        self.fence: int | None = None

    async def acquire(self) -> bool:
        token = str(uuid.uuid4())
        if not await self.r.set(self.key, token, nx=True, px=self.ttl):
            return False
        self.token = token
        self.fence = await self.r.incr(self.key + ":fence")
        return True

NX makes the set conditional on the key not existing, and PX attaches the lease in the same atomic command — never SET then EXPIRE separately, or a crash between them leaves a lock that never expires. The value is a random token unique to this acquisition; it is how release knows the lock is still yours. The fence is a separate counter incremented on every successful acquisition, so it only ever grows.

Verify: two concurrent acquire() calls on the same name — exactly one returns True.

The three parts of a safe Redis lock A flow of 4 stages. The three parts of a safe Redis lock SET key token NX PX atomic acquire + lease INCR key:fence monotonic fencing token write(..., fence) storage rejects stale compare-and-delete release only your own The lease limits how long a dead holder blocks others; the fence limits what a living stale holder can do.

2. Release only your own lock

Unconditional DEL removes whatever lock is there — possibly one that a different worker acquired after your lease expired. Compare and delete atomically in a Lua script:

RELEASE = """
if redis.call('get', KEYS[1]) == ARGV[1] then
  return redis.call('del', KEYS[1])
else
  return 0
end
"""


async def release(self) -> bool:
    return bool(await self.r.eval(RELEASE, 1, self.key, self.token))

Lua scripts run atomically in Redis, so nothing can change the key between the comparison and the delete. Verified: after worker A's lease expired and worker B acquired the lock, A's release returned 0 and B's lock was still there. Without the script, A's release would have removed B's lock, and a third worker could then acquire it while B was still working — two holders, created by a cleanup step.

Verify: release from a holder whose lease has expired returns False and leaves the current holder's key untouched.

3. Fence every write to the protected resource

Release and lease handling cannot stop a holder that does not know it lost the lock. Only the protected resource can, if every write carries the fence and the resource refuses fences older than the newest it has seen:

FENCED_WRITE = """
local current = tonumber(redis.call('hget', KEYS[1], 'fence') or '0')
if tonumber(ARGV[1]) < current then return 0 end
redis.call('hset', KEYS[1], 'fence', ARGV[1], 'value', ARGV[2])
return 1
"""


async def fenced_write(r, resource: str, fence: int, value: str) -> bool:
    return bool(await r.eval(FENCED_WRITE, 1, resource, fence, value))

For a Postgres table, the same check is a conditional update:

async def fenced_update(conn, invoice_id: int, fence: int, status: str) -> bool:
    result = await conn.execute(
        "update invoices set status=$1, lock_fence=$2 where id=$3 and lock_fence <= $2",
        status, fence, invoice_id)
    return result.endswith(" 1")                    # 'UPDATE 1' when accepted

Measured: B wrote with fence 2 and was accepted; A, waking from its pause, wrote with fence 1 and was rejected. Any resource that supports a conditional write — a version column, an object store's If-Match, a compare-and-set — can enforce the fence.

Verify: simulate a pause longer than the lease in a test; the stale holder's write is rejected and the current holder's value remains.

The fence stops a stale holder A sequence of 6 messages between 4 participants. The fence stops a stale holder worker A Redis lock worker B storage acquire: fence 1 pause 400 ms > 300 ms lease acquire: fence 2 write fence 2: accepted write fence 1: rejected release: not mine, no-op Measured against Redis 7.4: token 1 was refused after token 2 had been accepted.

4. Pick the TTL from pause behaviour, and renew

The lease must outlast normal pauses but not delay failover unreasonably. Measure the pauses you actually have — event loop lag, GC pauses — and set the TTL several times the worst; then renew while working:

EXTEND = """
if redis.call('get', KEYS[1]) == ARGV[1] then
  return redis.call('pexpire', KEYS[1], ARGV[2])
else
  return 0
end
"""


async def extend(self) -> bool:
    return bool(await self.r.eval(EXTEND, 1, self.key, self.token, self.ttl))

Extension, like release, must check ownership: extending a lock that now belongs to someone else would steal time from them. A renewal that returns False means the lease is gone; stop the work, as described in renewing lock leases with a heartbeat task. Keep fence unchanged on renewal — it identifies the acquisition, not the renewal.

Verify: a holder that renews every third of the TTL keeps the lock across a run several times longer than the TTL.

5. Use it as an async context manager

Wrap acquisition, waiting and release so call sites cannot get them wrong:

import asyncio
import random
from contextlib import asynccontextmanager


@asynccontextmanager
async def redis_lock(r, name: str, ttl_ms: int = 10_000, wait: float = 5.0):
    lock = RedisLock(r, name, ttl_ms)
    deadline = asyncio.get_running_loop().time() + wait
    while not await lock.acquire():
        if asyncio.get_running_loop().time() > deadline:
            raise TimeoutError(f"could not acquire {name}")
        await asyncio.sleep(random.uniform(0.05, 0.15))      # jitter avoids lockstep retries
    try:
        yield lock                                          # lock.fence goes into every write
    finally:
        await lock.release()


async def close_invoice(r, db, invoice_id: int) -> None:
    async with redis_lock(r, f"invoice:{invoice_id}") as lock:
        async with db.acquire() as conn:
            if not await fenced_update(conn, invoice_id, lock.fence, "closed"):
                raise RuntimeError("lost the lock before writing")

Polling with jitter is simple and adequate for coarse-grained locks. For high-contention locks, a Redis pub/sub notification on release or a queue is better than many pollers. Choose between this and a database lock with using Postgres advisory locks from asyncio.

Verify: two tasks contending for one lock run their critical sections strictly one after the other.

Common Redis lock mistakes and their fixes A grid of 4 rows by 3 columns. Common Redis lock mistakes and their fixes mistake what goes wrong fix SET, then EXPIRE separately crash leaves a lock forever SET NX PX in one command DEL to release deletes another holder's lock compare-and-delete script no fencing token stale holder overwrites newer work fence checked at the resource PEXPIRE without a check extends someone else's lease compare-and-extend script Each fix is a few lines; each missing one has caused real double-processing incidents.

Verification

The lock is safe when:

  • Acquisition is one atomic SET NX PX with a unique token per acquisition.
  • Release and extension compare the token inside a Lua script.
  • Every protected write carries the fence, and the resource rejects stale fences.
  • Tests simulate a pause longer than the lease and confirm the stale write is rejected.

Diagnostic Hook: count rejected fenced writes and lock acquisitions per name. Any rejection means a holder outlived its lease — investigate the pause behind it, usually event loop lag or a slow dependency inside the critical section. A high acquisition-failure rate on a name means heavy contention; consider whether the work really needs global exclusion or could be partitioned.

Pitfalls & edge cases

  • Redis failover. A replica promoted before replicating the lock key lets a second holder acquire; fencing makes it harmless.
  • Clock-based reasoning. Do not compute "my lease still has 200 ms" from local clocks; check with Redis or rely on the fence.
  • Fence counter without expiry. It must persist as long as resources remember fences; never let it reset to zero.
  • Very short TTLs. Below typical loop lag, holders lose locks constantly.

Frequently Asked Questions

How do I implement a distributed lock with Redis and asyncio?

Acquire with SET key token NX PX ttl using redis.asyncio, release with a Lua script that deletes the key only if it still holds your token, renew with a similar compare-and-extend script, and attach a fencing token from INCR to every protected write.

What is a fencing token in a Redis lock?

A number from a counter incremented on every acquisition. The protected resource rejects writes whose token is lower than one it has already accepted, so a holder that paused past its lease cannot overwrite newer work. In testing, token 1 was rejected after token 2.

Why should I not release a Redis lock with DEL?

Because if your lease expired and another worker acquired the lock, DEL removes their lock. Compare the stored token and delete atomically in a Lua script.

Is a single Redis node safe for distributed locks?

Safe enough when writes are fenced or idempotent, because a lock lost in a failover then cannot cause damage. If correctness depends on the lock alone, use a consensus store or let the resource enforce exclusivity.