Racing Coroutines and Taking the First Result¶
Racing is the pattern where you start the same logical operation against several sources — replicas, mirrors, regions, cache tiers — and take whichever answers first. It is how you cut tail latency when one replica is having a slow moment, and how you fall back without paying for sequential timeouts. The naive implementation has two bugs that show up in production: it returns the first task to finish, which is often the first to fail, and it leaves the losers running. In a test racing three replicas where the fastest one failed at 30 ms, a correct first_success() returned the second replica's answer at 80 ms and cancelled the third; the naive version raised the 30 ms ConnectionError and left two requests in flight.
Prerequisites¶
- Python 3.11+ (uses
asyncio.timeoutandExceptionGroup), stdlib only. asyncio.waitsemantics, from using asyncio.wait with FIRST_COMPLETED.- Cancellation basics, from Cancellation Patterns.
1. Write first_success, not first_completed¶
The building block keeps waiting until a task succeeds, collects every failure along the way, and cancels everything still running on the way out:
import asyncio
from collections.abc import Coroutine
from typing import Any, TypeVar
T = TypeVar("T")
async def first_success(*coros: Coroutine[Any, Any, T], timeout: float | None = None) -> T:
tasks = [asyncio.create_task(c) for c in coros]
errors: list[BaseException] = []
try:
async with asyncio.timeout(timeout):
pending = set(tasks)
while pending:
done, pending = await asyncio.wait(pending, return_when=asyncio.FIRST_COMPLETED)
for t in done:
if t.exception() is None:
return t.result()
errors.append(t.exception())
raise ExceptionGroup("all racers failed", errors)
finally:
for t in tasks:
t.cancel()
await asyncio.gather(*tasks, return_exceptions=True)
Three things make it correct. The loop continues past failures, so a fast error does not win. The finally cancels and awaits every task, whether the function returns, raises, times out or is itself cancelled — losers never outlive the race. And when every racer fails, the caller gets all the errors as an ExceptionGroup instead of whichever happened to be last, which is what you need to tell "all replicas down" from "one replica down".
Verify: race a failing fast coroutine against a slower succeeding one; you should get the slower one's result, and the cancelled set should contain every other racer still running.
2. Make losers cancel cleanly¶
Cancelling a loser only helps if the loser actually stops. A racer that does its I/O through an async client stops at its current await, and the client closes or returns the connection. A racer whose work is in a thread does not stop — asyncio.to_thread cancellation abandons the await but the thread runs to completion.
async def query_replica(pool, sql: str):
async with pool.acquire() as conn: # released on cancellation
try:
return await conn.fetch(sql)
except asyncio.CancelledError:
log.debug("lost the race on %s", pool.name)
raise # always re-raise
For database racers the cancellation should also cancel the query, not just the client-side wait; asyncpg sends a cancel request to the server when a fetch is cancelled, and other drivers differ. For HTTP racers, cancellation closes the stream mid-response, which means the connection is discarded rather than returned to the pool — a cost of racing worth measuring under load.
Verify: after each race, the number of in-flight requests to each replica drops back to zero; query the server's own view (active queries, open streams), not just your client's.
3. Stagger the starts: hedging¶
Firing every racer at once doubles or triples load on the backends for every request. Hedging fires the first racer immediately and the next only if the first has not answered within a delay, usually the backend's p95:
async def hedged(make_call, replicas: list[str], hedge_after: float):
tasks: list[asyncio.Task] = []
try:
for i, replica in enumerate(replicas):
tasks.append(asyncio.create_task(make_call(replica)))
if i == len(replicas) - 1:
break
done, _ = await asyncio.wait(tasks, timeout=hedge_after,
return_when=asyncio.FIRST_COMPLETED)
for t in done:
if t.exception() is None:
return t.result()
return await first_success(*(_await(t) for t in tasks))
finally:
for t in tasks:
t.cancel()
await asyncio.gather(*tasks, return_exceptions=True)
async def _await(task: asyncio.Task):
return await task
With hedge_after at the p95, only about 5% of requests send a second call, so the extra load is small while the slow tail is cut. The full treatment, including budgets that cap hedging under overload, is in hedging requests to cut tail latency.
Verify: count backend calls per request; with hedging at p95 it should average about 1.05, not 2 or 3.
4. Race against a timeout or a signal¶
Racing is not only between equivalent sources. Racing work against a shutdown event or a user cancellation is the same shape, and first_success is the wrong helper because the signal "succeeding" means the work did not:
async def unless_stopped(work, stop: asyncio.Event):
work_task = asyncio.create_task(work)
stop_task = asyncio.create_task(stop.wait())
try:
done, _ = await asyncio.wait({work_task, stop_task},
return_when=asyncio.FIRST_COMPLETED)
if work_task in done:
return work_task.result()
raise asyncio.CancelledError("stopped")
finally:
for t in (work_task, stop_task):
t.cancel()
await asyncio.gather(work_task, stop_task, return_exceptions=True)
For a plain time limit, prefer asyncio.timeout() — it cancels the work at the deadline with no extra task. The event race earns its place when the stop condition is not a time: a client disconnect, a shutdown signal, or a newer request superseding this one (the "latest request wins" pattern for search-as-you-type).
Verify: set the event while the work is running; the work task must be cancelled and the function must exit within one loop iteration.
5. Record who won¶
Racing hides problems by design: if replica A fails 30% of the time but B always answers, users never notice — until B is down too. Record the winner and every loser's outcome:
async def first_success_observed(named: dict[str, Coroutine], metrics) -> Any:
async def tagged(name, coro):
try:
result = await coro
metrics.inc("race_outcome", replica=name, outcome="success")
return name, result
except asyncio.CancelledError:
metrics.inc("race_outcome", replica=name, outcome="cancelled")
raise
except Exception:
metrics.inc("race_outcome", replica=name, outcome="error")
raise
winner, result = await first_success(*(tagged(n, c) for n, c in named.items()))
metrics.inc("race_winner", replica=winner)
return result
A replica whose error rate climbs while it never wins is a silent failure the race is masking; alert on per-replica error rate, not on the race's overall success rate.
Verify: inject failures into one replica; the race keeps succeeding and the per-replica error metric rises.
Verification¶
A race is implemented correctly when:
- A fast failure never wins: the returned value always comes from a successful racer.
- No racer outlives the race, including when the caller is cancelled or times out.
- All-failed produces every error, as an
ExceptionGroup, not just the last one. - Backend load matches the strategy: about N calls when racing all, about 1.05 when hedging at p95.
Diagnostic Hook: export a winner counter and an outcome counter per replica, and a histogram of time-to-first-success. A winner distribution that shifts sharply toward one replica is an early warning that the others are degrading; a time-to-first-success that tracks the slowest replica means the racers are not actually independent — usually a shared pool or a shared lock.
Pitfalls & edge cases¶
- Racers sharing one connection pool. If all racers wait on the same pool, they are not independent; a saturated pool makes every racer slow at once.
- Racing non-idempotent operations. Two racers that both write can both succeed; only race reads or operations protected by an idempotency key.
- Thread-backed racers. Cancellation does not stop the thread, so losers keep consuming resources until they finish.
as_completed()for racing. It does not cancel the rest when you break out of the loop; you still need thefinally.- Swallowed
CancelledErrorin a racer. The cleanupgatherthen waits for that racer to finish naturally, and the race takes as long as the slowest replica.
Frequently Asked Questions¶
How do I run several coroutines and return the first result in asyncio?
Wrap each in a task, loop on asyncio.wait with FIRST_COMPLETED, return the first task whose exception() is None, and cancel and await every other task in a finally block. Do not return the first task to finish without checking it succeeded.
Does asyncio have a built-in first-completed race?
asyncio.wait with return_when=FIRST_COMPLETED tells you which tasks finished first, but it does not check for success and does not cancel the others. A small first_success helper adds both.
What happens to the losing coroutines?
Nothing, unless you cancel them. They keep running and holding connections. Cancel them and await the cancellation so they release resources before the race returns.
Is racing the same as hedging?
Racing starts every attempt at once; hedging starts the next attempt only if the previous one has not answered within a delay such as the p95. Hedging gives most of the tail-latency benefit for a small amount of extra load.
Related¶
- Coroutine Design Patterns — up to the topic overview.
- Processing results in completion order with as_completed — when you want every result, not just the first.
- Asyncio Fundamentals & Event Loop Architecture — the section overview.