Hashing Passwords Without Blocking the Event Loop¶
Password hashing is deliberately slow: bcrypt, scrypt and Argon2 are designed to cost tens or hundreds of milliseconds per attempt so that stolen hashes are expensive to crack. In an async service that cost lands on the event loop unless it is moved, and every request on the worker waits while a login is checked. Measured on Python 3.14 on a 24-core machine: one bcrypt.checkpw at cost factor 12 took 171 ms, and one Argon2id verify with argon2-cffi's defaults took 122 ms. Thirty-two concurrent logins verified on the event loop ran at 5.9 logins per second (bcrypt) and 8.6 (Argon2id), and loop lag reached 5,444 ms and 3,723 ms — every other request on the worker waited that long. With asyncio.to_thread, bcrypt reached 79.7 logins per second with lag p99 of 21 ms, because it releases the GIL; Argon2id reached 17.9 with 5 ms lag. A process pool of eight gave 42.0 and 17.2. This guide verifies passwords without stalling everything else.
Prerequisites¶
- Python 3.11+,
pip install bcrypt argon2-cffi. - Offloading CPU work, from offloading CPU work with loop.run_in_executor.
- The topic overview, Securing Async Services.
1. Measure what a hash costs on the loop¶
Time one verification, then run concurrent logins on the loop with a lag probe alongside:
import asyncio, time
import bcrypt
from argon2 import PasswordHasher
PW = b"correct horse battery staple"
BCRYPT_HASH = bcrypt.hashpw(PW, bcrypt.gensalt(rounds=12))
PH = PasswordHasher() # argon2id with library defaults
ARGON_HASH = PH.hash(PW.decode())
async def login_on_loop() -> bool:
return bcrypt.checkpw(PW, BCRYPT_HASH) # 171 ms of CPU on the event loop thread
Measured: 32 concurrent logins verified this way took about 5.4 s in total at 5.9 logins per second, and the event-loop lag probe recorded a single stall of 5,444 ms — the loop executed the 32 hashes back to back with nothing else running in between. For Argon2id, 8.6 logins per second and a 3,723 ms stall. Under real traffic that means health checks time out and every unrelated request on the worker waits for the burst of logins to finish, exactly the symptom described in finding blocking calls with asyncio debug mode.
Verify: with asyncio debug mode enabled in staging, the login handler does not appear in "Executing ... took" warnings.
2. Move verification to a thread¶
Both libraries do their work in C and release the GIL while hashing, so a thread is enough to keep the loop free:
async def verify_password(password: str, stored_hash: bytes) -> bool:
return await asyncio.to_thread(bcrypt.checkpw, password.encode(), stored_hash)
async def verify_password_argon2(password: str, stored_hash: str) -> bool:
try:
return await asyncio.to_thread(PH.verify, stored_hash, password)
except argon2.exceptions.VerifyMismatchError:
return False
Measured: bcrypt in to_thread reached 79.7 logins per second — the default pool's 28 threads hashed in parallel across the 24 cores — and loop lag p99 stayed at 21 ms. Argon2id reached 17.9 logins per second with 5 ms lag; its default parameters already use several lanes and 64 MiB of memory per hash, so it saturates memory bandwidth sooner and parallelizes less. Either way the event loop kept serving other requests while logins were checked. The thread approach depends on the library releasing the GIL; libraries that hash in pure Python would not benefit, which is why the measurement, not the assumption, should decide.
Verify: loop lag during a burst of logins stays within your latency budget.
3. Bound how many hashes run at once¶
Offloading moves the cost off the loop; it does not limit it. A burst of login attempts — legitimate or a credential-stuffing attack — can fill the default executor, delaying every other to_thread call and even DNS lookups, as shown in resolving DNS without blocking the executor. Give hashing its own bounded pool:
from concurrent.futures import ThreadPoolExecutor
HASH_POOL = ThreadPoolExecutor(max_workers=8, thread_name_prefix="pwhash")
HASH_SLOTS = asyncio.Semaphore(32) # attempts waiting or running
async def verify_password(password: str, stored_hash: bytes) -> bool:
if HASH_SLOTS.locked():
raise TooManyAttempts() # 429/503 instead of an unbounded queue
async with HASH_SLOTS:
loop = asyncio.get_running_loop()
return await loop.run_in_executor(HASH_POOL, bcrypt.checkpw, password.encode(), stored_hash)
A dedicated pool caps the CPU that hashing can take — eight workers here — and isolates it from other blocking work. The semaphore bounds waiting attempts, so an attack produces fast rejections rather than a growing queue; combine it with per-account and per-IP rate limits, as in per-tenant rate limits in async services. Size the pool from the CPU budget you are willing to spend on authentication.
Verify: under a simulated credential-stuffing burst, other endpoints' latency is unaffected and excess login attempts are rejected quickly.
4. Choose a process pool for hashes that hold the GIL¶
If a hashing implementation does not release the GIL — a pure-Python KDF, or a library compiled without that support — threads give no parallelism and still contend with the event loop for the interpreter. A process pool sidesteps both:
from concurrent.futures import ProcessPoolExecutor
HASH_PROCS = ProcessPoolExecutor(max_workers=4)
async def verify_password(password: str, stored_hash: bytes) -> bool:
loop = asyncio.get_running_loop()
return await loop.run_in_executor(HASH_PROCS, bcrypt.checkpw, password.encode(), stored_hash)
Measured with bcrypt: an 8-process pool reached 42.0 logins per second with loop lag p99 of 1.1 ms — the lowest lag of any option, because nothing ran in the service's own process — but lower throughput than 28 threads, simply because it had fewer workers. Each call pickles the password and hash to the worker and the result back, which is negligible here. On Python 3.14, the default process start method on Linux is forkserver, so the module that creates the pool needs an if __name__ == "__main__": guard when run as a script; without it the measurement script itself failed to start its workers.
Verify: for your hashing library, compare loop lag and throughput between to_thread and a process pool under the same burst, and pick by measurement.
5. Keep hashing parameters honest¶
Moving hashes off the loop removes the temptation to weaken them for latency's sake. Choose parameters by target duration on production hardware, and rehash on login when they change:
PH = PasswordHasher(time_cost=3, memory_cost=64 * 1024, parallelism=4) # argon2-cffi defaults
async def login(username: str, password: str) -> User | None:
user = await users.get(username)
if user is None:
await asyncio.to_thread(PH.hash, password) # same cost for unknown users
return None
try:
await asyncio.to_thread(PH.verify, user.password_hash, password)
except argon2.exceptions.VerifyMismatchError:
return None
if PH.check_needs_rehash(user.password_hash):
new_hash = await asyncio.to_thread(PH.hash, password)
await users.update_hash(user.id, new_hash)
return user
Hashing a dummy password for unknown usernames keeps response times similar whether or not the account exists, so timing does not reveal which usernames are registered. check_needs_rehash upgrades stored hashes to new parameters as users log in. With the work off the loop, a 100–200 ms hash is a per-login cost, not a per-request outage.
Verify: response times for unknown and known usernames with wrong passwords are indistinguishable in a test.
Verification¶
Password hashing is safe for an async service when:
- No hash runs on the event loop thread.
- Hashing has its own bounded pool, and excess attempts are rejected quickly.
- The pool type matches the library: threads when it releases the GIL, processes when it does not.
- Parameters stay strong, unknown users cost the same as known ones, and hashes are upgraded on login.
Diagnostic Hook: chart login attempts per second, hash-pool queue depth and event-loop lag together. A rising queue with flat lag is the pool doing its job under load; rising lag with logins means a hash is still running on the loop somewhere — often in a code path that was missed, such as password changes or account creation.
Pitfalls & edge cases¶
- Hashing on the loop. Measured: a 5.4 s stall for 32 concurrent bcrypt logins.
- The default executor for hashing. A login burst then delays every other
to_threadcall. - Weakening cost factors for latency. Offload instead; keep the hash slow.
- Forgetting other hash paths. Registration and password changes hash too.
Frequently Asked Questions¶
Does bcrypt block the asyncio event loop?
Yes, if called directly in a coroutine: one checkpw at cost 12 took 171 ms, and 32 concurrent logins stalled the loop for 5.4 s in testing. Run it with asyncio.to_thread or in an executor.
Is asyncio.to_thread enough for password hashing?
For bcrypt and argon2-cffi, which release the GIL, yes: bcrypt reached 79.7 logins per second with loop lag p99 of 21 ms. Use a dedicated, bounded pool so login bursts cannot fill the default executor.
Should I use a process pool for password hashing?
When the implementation holds the GIL, or to minimize loop lag: an 8-process pool kept lag p99 at 1.1 ms. For GIL-releasing libraries, threads gave higher throughput in testing.
How do I prevent login timing attacks in an async service?
Hash a dummy password when the username does not exist, so the response takes as long as a real verification, and run both off the event loop.
Related¶
- Securing Async Services — up to the topic overview.
- Preventing regex denial of service in async handlers — CPU work that a thread does not fix.
- Resilience, Cancellation & Error Handling — the section overview.