Choosing Between Redis and Postgres for Coordination¶
Most asyncio services that need a distributed lock already run Redis, Postgres or both, and the choice between them is usually made by habit. They fail differently, and that difference matters more than speed. Measured with Redis 8.10 and Postgres 17.11 in local containers, redis-py 5.3.1 and asyncpg 0.31.0: an uncontended lock and unlock took 181 µs at the median with Redis (SET NX PX plus a Lua release) and 105 µs with a Postgres session advisory lock. With 50 contenders on one key, Postgres's server-side wait queue handed the lock over 8,349 times a second against 1,956 for polling Redis clients. When a holder was killed with SIGKILL, the next acquirer got the Redis lock after 4.82 s — the rest of the 5 s lease — and the Postgres lock after 0.00 s, because the dead connection released it. When a holder was stopped for 8 s, Redis handed the lock to someone else after 4.84 s while the stalled holder still believed it held it, and Postgres never handed it over at all — the waiter was still blocked minutes later — until idle_session_timeout turned the session into a lease and released it in 1.77 s. With the lock and the data in the same Postgres transaction, six 0.8 s stalls cost zero lost updates without any fencing token. This guide shows how to choose.
Prerequisites¶
- Redis and/or Postgres reachable from your services; redis-py 5+ and asyncpg.
- Both lock mechanisms, from implementing a Redis lock with fencing tokens and using Postgres advisory locks from asyncio.
- The topic overview, Distributed Locks & Coordination.
1. Compare the cost of a lock round trip¶
Measure the basic operation first, uncontended and on the same machine as your service would be:
# Redis: acquire with a lease, release only if still ours
token = uuid.uuid4().hex
await r.set("lock:report", token, nx=True, px=5000)
await release(keys=["lock:report"], args=[token]) # Lua compare-and-delete
# Postgres: session-level advisory lock
await conn.fetchval("SELECT pg_try_advisory_lock(42)")
await conn.fetchval("SELECT pg_advisory_unlock(42)")
# Postgres: transaction-level advisory lock, released by COMMIT
async with conn.transaction():
await conn.fetchval("SELECT pg_try_advisory_xact_lock(42)")
Measured over 3,000 iterations each: Redis p50 181 µs, p99 333 µs; Postgres session lock p50 105 µs, p99 214 µs; Postgres transaction lock including BEGIN and COMMIT p50 170 µs, p99 321 µs. All three are two network round trips, and the client libraries' overhead dominated. Speed is rarely the deciding factor: a lock that guards work taking milliseconds or more costs a negligible fraction of it with either store.
Verify: you know the lock overhead relative to the work it protects for your service.
2. Compare behaviour under contention¶
Contention is where the designs diverge. A Redis lock has no waiting list: a client that fails SET NX must poll. A Postgres advisory lock queues waiters in the server and wakes the next one the moment the lock is released:
# Redis: poll until free
while not await r.set("lock:hot", token, nx=True, px=5000):
await asyncio.sleep(0.001)
# Postgres: block in the server until granted
await conn.execute("SELECT pg_advisory_lock(43)")
Measured with 50 asyncio contenders taking the lock 20 times each: Postgres completed 8,349 handoffs per second, Redis 1,956 with a 1 ms poll. Polling also makes Redis unfair — whoever polls at the right moment wins — as measured in distributed semaphores with Redis. The price on the Postgres side is connections: a waiting or holding session occupies a database connection for its whole duration, so 50 contenders held 50 connections. With a pool sized for queries, a burst of lock waiters can starve ordinary queries of connections.
Verify: under your expected contention, lock waits do not exhaust the database connection pool.
3. Compare what happens when a holder dies or stalls¶
A Redis lock is a lease: it expires on a timer whatever the holder is doing. A Postgres advisory lock belongs to a session: it lasts exactly as long as the database connection. That gives opposite behaviour in the two common failures:
# Postgres: make a stalled session time out, so the lock behaves like a lease
conn = await asyncpg.connect(dsn, server_settings={"idle_session_timeout": "2000"})
await conn.execute("SELECT pg_advisory_lock(7)")
Measured: a holder killed with SIGKILL released its Postgres lock immediately — the kernel closed its socket — while the Redis lock stayed held for the remaining 4.82 s of its lease. A holder stopped with SIGSTOP for 8 s kept its Postgres lock for the whole stall and beyond: after resuming it went on holding it, and the waiter was still blocked when the test was stopped minutes later. Redis gave the lock to the waiter after 4.84 s, and the stalled holder, on waking, reported every second that it "still thinks it holds" while the key held the new owner's value. Setting idle_session_timeout to 2 s on the holder's connection made Postgres terminate the stalled session: the waiter acquired after 1.77 s and the holder, on resuming, got ConnectionDoesNotExistError — it learned it had lost the lock, which a Redis holder never does on its own. Such a timeout also ends healthy sessions that stay idle while holding a lock, so a long holder must issue a cheap query periodically.
Verify: for each lock, you know whether a stalled holder blocks others forever (session) or can act after losing the lock (lease).
4. Keep the lock and the data in one place when you can¶
When the data the lock protects lives in Postgres, a transaction-level advisory lock gives a property no external lock can: the lock and the writes commit or roll back together. A stalled holder's session is terminated, its transaction aborted, and its writes never land:
conn = await asyncpg.connect(dsn, server_settings={"idle_in_transaction_session_timeout": "500"})
async with conn.transaction():
await conn.execute("SELECT pg_advisory_xact_lock(99)")
value = await conn.fetchval("SELECT value FROM counter WHERE id = 1")
await conn.execute("UPDATE counter SET value = $1 WHERE id = 1", value + 1)
Measured with three processes for 10 s and six random 0.8 s SIGSTOP pauses: 284 successful writes, final counter 284, and 6 transactions aborted by the server's timeout — zero lost updates, without a fencing token. The same pause pattern against a Redis lock protecting the same Postgres counter lost 3–4 updates per run, as measured in testing distributed locks under pauses, and needed a fence column to fix. When the protected resource is elsewhere — an object store, an external API, files — no transaction spans it, and both options need fencing or idempotency at that resource.
Verify: locks protecting Postgres data are transaction-scoped in the same database, with an in-transaction idle timeout.
5. Decide by failure mode and by where the data lives¶
Put the measurements together into a choice:
def choose_lock(data_in_postgres: bool, holders_may_stall: bool, many_waiters: bool) -> str:
if data_in_postgres:
return "pg_advisory_xact_lock in the same transaction, idle_in_transaction timeout"
if many_waiters:
return "Postgres session lock (server queue) with idle_session_timeout, or Redis with notify"
if holders_may_stall:
return "Redis lease + fencing token checked by the protected resource"
return "either; prefer the store you already operate"
Two operational differences finish the comparison. A Redis instance that crashes loses locks that were not yet persisted: a key set just before docker kill was gone after restart, so a holder carried on while a new holder could acquire the same lock — while a graceful restart, with the default snapshot-on-shutdown configuration, preserved the key and its remaining TTL. A Postgres restart (1.3 s here) ended every session: holders got InterfaceError on their next query, and the next acquirer succeeded immediately. And Postgres locks cost connections: each holder and waiter needs one, which matters with a small pool or behind PgBouncer in transaction mode, where session-level advisory locks do not work reliably and only the transaction-level form should be used.
Verify: each lock in your system has a documented choice and the failure behaviour it accepts.
Verification¶
The coordination store fits the lock when:
- Lock overhead is small relative to the protected work, in either store.
- Stalled holders are handled: leases with fencing in Redis, idle timeouts in Postgres.
- Locks protecting Postgres data share its transaction.
- Crash and restart behaviour of the store has been tested, not assumed.
Diagnostic Hook: for Postgres, query pg_locks joined with pg_stat_activity for locktype = 'advisory' and alert on any lock held by a session whose state_change is older than your longest legitimate hold — that is a stalled holder blocking everyone. For Redis, alert on fence rejections at the protected resource, which count the times a lease expired under a live holder.
Pitfalls & edge cases¶
- Postgres session locks with stalled holders. Measured: held indefinitely without
idle_session_timeout. - Redis locks after a crash. Unsaved keys vanished on
docker kill. - Session advisory locks behind PgBouncer transaction pooling. Use transaction-level locks.
- Connection pools sized only for queries. Lock waiters occupy connections too.
Frequently Asked Questions¶
Should I use Redis or Postgres for distributed locks?
If the protected data is in Postgres, a transaction-level advisory lock in the same transaction lost no updates under stalls without fencing. For external resources, a Redis lease with fencing tokens is simpler to bound in time; Postgres session locks are held as long as the connection lives.
Which is faster, Redis locks or Postgres advisory locks?
In testing, Postgres was slightly faster uncontended (105 µs vs 181 µs median) and much faster under contention (8,349 vs 1,956 handoffs/s), because waiters queue in the server instead of polling.
What happens to a Postgres advisory lock if the client crashes?
It is released when the connection closes; a SIGKILLed holder's lock was free immediately. A stalled but connected client keeps it, unless idle_session_timeout or idle_in_transaction_session_timeout ends the session.
Are Redis locks lost when Redis restarts?
After a crash, keys not yet persisted are lost; in testing a lock set just before docker kill was gone after restart. A graceful restart with snapshot-on-shutdown kept it.
Related¶
- Distributed Locks & Coordination — up to the topic overview.
- Deduplicating work across replicas — when idempotency beats a lock.
- Concurrent Execution & Worker Patterns — the section overview.