Skip to content

Implementing Per-Key Async Locks

A global asyncio.Lock around "update this user's balance" is correct and throws away almost all concurrency: two different users wait for each other. What you want is one lock per key — per user, per order id, per file path — so operations on the same key are serialised and operations on different keys run in parallel. In a test with 1,000 operations of 1 ms spread over 50 keys, a global lock took 1.08 s; a per-key lock took 0.03 s, with zero overlapping operations on any key. The difficulty is not the locking but the bookkeeping: a naive dict[key, Lock] grows by one entry for every key ever seen, which in a long-running service is a memory leak keyed by your user base. The version below removed every lock once its last waiter left — 0 entries remained after the run.

Prerequisites

1. Count waiters, delete on the last release

The lock table needs a reference count: how many tasks currently hold or wait for each key's lock. When it drops to zero, nobody can be relying on that lock object, and the entry can go:

import asyncio
from collections.abc import AsyncIterator, Hashable
from contextlib import asynccontextmanager


class KeyedLock:
    def __init__(self) -> None:
        self._locks: dict[Hashable, asyncio.Lock] = {}
        self._users: dict[Hashable, int] = {}

    @asynccontextmanager
    async def lock(self, key: Hashable) -> AsyncIterator[None]:
        lk = self._locks.get(key)
        if lk is None:
            lk = self._locks[key] = asyncio.Lock()
        self._users[key] = self._users.get(key, 0) + 1
        try:
            async with lk:
                yield
        finally:
            self._users[key] -= 1
            if self._users[key] == 0:
                del self._users[key]
                del self._locks[key]

    def __len__(self) -> int:
        return len(self._locks)

The get-or-create and the increment happen with no await between them, so on a single event loop they are atomic: no other task can observe a lock with a zero count that is about to be used. The decrement is in finally, so a cancelled waiter — one that gave up while queued for the lock — also releases its reference. Without that, every timeout would leave a permanent entry.

Verify: after a burst of operations, len(keyed) == 0.

1,000 one-millisecond operations over 50 keys 2 horizontal bars comparing one global lock with the others. 1,000 one-millisecond operations over 50 keys one global lock 1.08 s one lock per key 0.03 s, 0 overlaps per key Each operation holds its lock for 1 ms; 20 operations per key. Per-key locks keep the same-key guarantee and give back the concurrency between keys.

2. Use it around the read-modify-write

The lock protects a sequence of awaits that must not interleave for the same key — typically read, compute, write against an external store:

locks = KeyedLock()


async def credit(user_id: int, amount: int) -> int:
    async with locks.lock(("balance", user_id)):
        balance = await db.fetchval("select balance from accounts where id=$1", user_id)
        new = balance + amount
        await db.execute("update accounts set balance=$1 where id=$2", new, user_id)
        return new

Namespacing the key — ("balance", user_id) rather than user_id — keeps unrelated features from serialising against each other by accident when they share an id space.

Be clear about what this protects against: concurrent tasks in this process. A second process or a second replica has its own lock table and will interleave freely. For a single-row update like this one, the database can do the work atomically (update … set balance = balance + $1) or with SELECT … FOR UPDATE, which is better than any in-process lock. Per-key locks shine where the critical section spans things the database cannot lock — a cache, a file, a call to an external API. Cross-process versions are in Distributed Locks & Coordination.

Verify: run 100 concurrent credits of 1 against one user; the final balance increases by exactly 100.

3. Bound how long callers wait

Under contention a hot key becomes a queue. Without a bound, callers pile up behind a slow operation and each holds a request open. Put the timeout around the acquisition:

async def credit_with_deadline(user_id: int, amount: int) -> int:
    try:
        async with asyncio.timeout(2):
            async with locks.lock(("balance", user_id)):
                return await apply_credit(user_id, amount)
    except TimeoutError:
        raise Busy(f"user {user_id} is busy, retry later") from None

The timeout covers both waiting for the lock and the work inside it, which is usually what a request-level deadline wants. Because the waiter count is decremented in finally, a timed-out waiter leaves the table consistent. If you want to bound only the wait, acquire with a timeout and run the body outside it — but then the body needs its own deadline, as described in propagating deadlines across async service calls.

Verify: hold one key's lock for 5 s; other callers for that key fail after 2 s and len(locks) returns to zero afterwards.

Three callers on one hot key 4 lanes over time. Three callers on one hot key caller 1 holds lock: slow operation caller 2 waiting runs caller 3 waiting timeout waiter count 1 2 3 2, then 1, then removed time → Waiters that give up still decrement the count, so the table empties once the key goes quiet.

4. Know the fairness you get

asyncio.Lock wakes waiters in FIFO order, and per-key locks inherit that: for one key, operations run in arrival order. Across keys there is no fairness at all — a key with a thousand queued operations does not slow down a key with one. That is usually exactly right. Two cases where it is not:

  • Ordering requirements beyond one process. FIFO is per-process; if a client sends two updates for the same key and they land on different workers, their order is undefined. Route by key — the approach in processing queue items in order per key.
  • Starvation by a hot key. If one key's queue grows without bound, the tasks waiting on it hold memory and requests. Cap waiters per key and fail fast:
    @asynccontextmanager
    async def lock(self, key, max_waiters: int = 100):
        if self._users.get(key, 0) >= max_waiters:
            raise Busy(f"too many operations queued for {key!r}")
        ...  # as before

Verify: fire 1,000 operations at one key with max_waiters=100; 900 fail immediately rather than queueing.

5. Avoid deadlocks when a task needs two keys

Transferring money from A to B needs both keys. Two concurrent transfers, A→B and B→A, each taking their source lock first, deadlock. The standard fix is a global acquisition order: always lock keys in sorted order.

from contextlib import AsyncExitStack


async def transfer(src: int, dst: int, amount: int) -> None:
    async with AsyncExitStack() as stack:
        for key in sorted({("balance", src), ("balance", dst)}):
            await stack.enter_async_context(locks.lock(key))
        await debit(src, amount)
        await credit_unlocked(dst, amount)

AsyncExitStack releases the locks in reverse order on any exit path, including cancellation part-way through acquiring them. Note the call to credit_unlocked — calling the locked credit() from inside would try to take the same key again, and asyncio.Lock is not reentrant: the task would wait for itself forever.

Verify: run 1,000 random transfers between 10 accounts concurrently; it completes, and the total balance is unchanged.

Is a per-key lock the right tool? A decision on Who else touches this key with 3 outcomes. Is a per-key lock the right tool? Who else touches this key? only this process KeyedLock cheap, FIFO per key a single DB row the database atomic update or FOR UPDATE other processes too a distributed lock Redis or Postgres An in-process lock only coordinates tasks on one event loop.

Verification

Per-key locking is correct when:

  • Same-key operations never overlap, checked by an in-section counter in tests.
  • Different keys run concurrently — throughput scales with the number of active keys.
  • The lock table empties when traffic stops, including after timeouts and cancellations.
  • Multi-key operations acquire in a fixed order and never call a locked function while holding its key.

Diagnostic Hook: export len(keyed_lock) as a gauge, plus a histogram of lock wait time and the maximum waiter count per key. A gauge that grows with uptime means an exit path skips the decrement. A long tail in wait time concentrated on a handful of keys identifies hot keys, which usually need a different design — batching their updates, or moving them to an atomic database operation.

Pitfalls & edge cases

  • A defaultdict(asyncio.Lock) with no cleanup. One permanent lock per key ever seen.
  • Deleting the lock while waiters exist. A new caller then creates a second lock for the same key, and two tasks hold "the" lock at once.
  • Using weakref.WeakValueDictionary for the table. It can drop a lock between a caller fetching it and acquiring it unless every holder keeps a strong reference — the counting version is simpler to reason about.
  • Reentrancy. asyncio.Lock is not reentrant; a task that re-enters the same key deadlocks with itself.

Frequently Asked Questions

How do I lock per key in asyncio?

Keep a dict of asyncio.Lock objects keyed by the thing you are serialising, plus a count of tasks using each lock. Create the lock on first use, increment the count before acquiring, and decrement in a finally block, deleting both entries when the count reaches zero.

Why does my dictionary of asyncio locks keep growing?

Because nothing removes a lock after its last use. Track how many tasks hold or wait for each key and delete the entry when that count drops to zero.

Do per-key asyncio locks work across processes?

No. They coordinate tasks on one event loop in one process. Use the database's own locking or a distributed lock such as a Redis lock or a Postgres advisory lock across processes.

How do I lock two keys without deadlocking?

Always acquire them in a consistent global order, such as sorted by key, using AsyncExitStack so they are released in reverse order on any exit path.