Yielding Control with asyncio.sleep(0)¶
A coroutine that never awaits never lets anything else run. That is the whole contract of cooperative scheduling, and it bites hardest in code that looks async — an async def that parses a large payload, walks a big dictionary or scores a few thousand items in a loop. In a test, a loop doing 40,000 small units of work took 128 ms, and a heartbeat task that wanted to wake every 5 ms was late by 128 ms: the full duration of the loop. Adding await asyncio.sleep(0) every 100 iterations cut the heartbeat's worst lateness to 1.2 ms and cost about 4% in throughput. This guide covers what that one line does inside the event loop, how to choose the yield interval, and the point at which yielding stops being the right answer.
Prerequisites¶
- Python 3.11+; every example is stdlib only. Timings below were taken on Python 3.14 on Linux.
- The ready queue model from Task Scheduling & Lifecycle: a task runs until it awaits something that is not ready, then the loop picks the next callback.
- A way to see lag — the heartbeat technique from measuring event loop lag in production is used throughout.
1. Understand what sleep(0) does¶
asyncio.sleep(0) is special-cased. It does not create a timer: it yields a bare None up through the task's __step, which tells the task machinery "reschedule me immediately". The task's next step goes to the back of the ready queue, behind every callback that became ready while it was running — I/O completions, timer expirations, other tasks' wake-ups.
import asyncio
async def chatty(name: str) -> None:
for i in range(3):
print(name, i)
await asyncio.sleep(0) # back of the ready queue, no timer created
async def main() -> None:
async with asyncio.TaskGroup() as tg:
tg.create_task(chatty("a"))
tg.create_task(chatty("b"))
asyncio.run(main())
# a 0, b 0, a 1, b 1, a 2, b 2 — strict round-robin
Without the sleep(0), the output is a 0, a 1, a 2, b 0, b 1, b 2: task a runs to completion before b gets a turn. With it, the two tasks interleave one step at a time, because each yield puts the yielding task behind the other.
Any non-zero delay goes through a different path: a TimerHandle on the heap, a selector timeout, and the minimum resolution of the clock. sleep(0.001) in a hot loop therefore costs far more than sleep(0) and adds real latency to the task doing the work.
Verify: run the snippet and confirm strict alternation; remove the await and confirm one task runs to completion first.
2. Measure the stall before you fix it¶
Do not guess whether a loop needs yields. Run a heartbeat next to it and read the worst-case lateness:
import asyncio
import time
async def heartbeat(stop: asyncio.Event, lags: list[float]) -> None:
while not stop.is_set():
start = time.perf_counter()
await asyncio.sleep(0.005)
lags.append(time.perf_counter() - start - 0.005)
def work(i: int) -> int: # stand-in for parsing / scoring / transforming
return sum(k * i for k in range(200))
async def run(yield_every: int, n: int = 40_000) -> tuple[float, float]:
stop, lags = asyncio.Event(), []
hb = asyncio.create_task(heartbeat(stop, lags))
await asyncio.sleep(0.01)
t0 = time.perf_counter()
for i in range(n):
work(i)
if yield_every and i % yield_every == 0:
await asyncio.sleep(0)
elapsed = time.perf_counter() - t0
stop.set()
await hb
return elapsed * 1000, max(lags) * 1000
The numbers from this harness, 40,000 iterations of about 3 µs each:
| Yield every | Loop total | Worst heartbeat lag |
|---|---|---|
| never | 127.5 ms | 127.6 ms |
| 1 iteration | 187.4 ms | 0.9 ms |
| 10 iterations | 143.8 ms | 0.9 ms |
| 100 iterations | 132.7 ms | 1.2 ms |
| 1,000 iterations | 135.2 ms | 11.5 ms |
Never yielding is the fastest way to finish the loop and the worst thing for everything else on the loop. Yielding on every iteration fixed the lag but added 47% to the loop's own runtime — at 1.6 µs per bare sleep(0), a yield costs about half of what one unit of work does.
Verify: your own numbers will differ; what matters is the shape — lag tracks the time between yields, and throughput loss tracks the number of yields.
3. Pick the interval from a time budget, not a count¶
An iteration count is only a proxy. The real knob is time between yields, and the right value is a fraction of your latency budget — if the service promises a p99 of 50 ms, a 10 ms stall is already a fifth of it. Yield on elapsed time and the loop stays correct when the work per item changes:
import asyncio
import time
async def cooperative(items, fn, budget_s: float = 0.002):
"""Apply fn to every item, yielding whenever budget_s of CPU time has passed."""
out = []
deadline = time.perf_counter() + budget_s
for item in items:
out.append(fn(item))
if time.perf_counter() >= deadline:
await asyncio.sleep(0)
deadline = time.perf_counter() + budget_s
return out
time.perf_counter() costs well under 100 ns, so checking it every iteration is cheap compared with a yield. A 2 ms budget bounds the stall any neighbour sees to roughly 2 ms plus one item, regardless of whether an item takes 3 µs or 300 µs.
Verify: run the heartbeat harness against cooperative() with items of very different cost; the worst lag should stay near the budget in every case.
4. Know when yielding is the wrong tool¶
Yielding fixes fairness, not cost. The CPU work still happens on the event loop thread, so the loop's total capacity for everything else is reduced by exactly the CPU time the work consumes. Three signs it is time to move the work off the loop:
- The work is large relative to the request rate. A 128 ms computation per request at 20 requests per second is 2.5 s of CPU per second — no amount of yielding fits that on one thread.
- The work calls into C that holds the GIL in one long call — a big
json.loads, a regex over megabytes,hashlibon a large buffer. There is no Python loop to put a yield in. - Latency of the work itself matters. Yielding makes the work take longer, by design.
import asyncio
from concurrent.futures import ProcessPoolExecutor
async def score_batch(pool: ProcessPoolExecutor, batch: list[int]) -> list[int]:
loop = asyncio.get_running_loop()
return await loop.run_in_executor(pool, _score_all, batch) # zero loop stall
def _score_all(batch: list[int]) -> list[int]:
return [sum(k * i for k in range(200)) for i in batch]
The decision between threads and processes for that work is covered in offloading CPU work with run_in_executor; for pure-Python CPU work on a GIL build, it is a process pool.
Verify: with the work in a process pool, the heartbeat's worst lag stays at its idle level no matter how long the batch takes.
5. Put the yield where cancellation can reach it¶
A task can only be cancelled at an await. A CPU loop with no awaits is uncancellable for its whole duration — task.cancel() sets a flag that is not acted on until the loop finishes. Each sleep(0) is also a cancellation point, which is a second reason to have them:
async def cancellable_transform(rows):
done = 0
try:
for i, row in enumerate(rows):
transform(row)
done += 1
if i % 100 == 0:
await asyncio.sleep(0) # CancelledError is raised here
except asyncio.CancelledError:
log.info("transform cancelled after %d rows", done)
raise # always re-raise
With a yield every 100 rows, a cancellation landed 0.03 ms after cancel() in the harness above instead of after the whole loop. The same reasoning applies to asyncio.timeout(): a timeout can only fire at an await, so a CPU loop without yields overruns its deadline by however long it runs. Implementing cooperative cancellation in CPU loops covers the pattern for work that has moved into threads, where sleep(0) is not available.
Verify: cancel the task 10 ms into the loop and confirm it finishes within a millisecond or two, with the partial count logged.
Verification¶
The change is working when:
- A heartbeat beside the loop reports worst lag close to your yield budget instead of the loop's full duration.
- Loop throughput drops by single-digit percent, not tens of percent — if it drops more, you are yielding too often.
- Cancellation and timeouts take effect within one yield interval.
- In debug mode (
PYTHONASYNCIODEBUG=1), the "Executingtook N seconds" warnings for this task disappear, since no single step exceeds slow_callback_durationany more.
Diagnostic Hook: export the heartbeat's worst lag per minute as a gauge and alert when it exceeds a quarter of your latency budget. Pair it with the slow-callback log from tracing slow callbacks: the log names the task, the gauge tells you how bad it is.
Pitfalls & edge cases¶
sleep(0)inside a lock. Yielding while holding anasyncio.Locklets other tasks run but not the ones waiting for that lock; if they are the ones you wanted to unblock, nothing improved.- Yielding from a tight loop over a generator that itself blocks. If each
next()does a blocking read, the yield happens between stalls rather than preventing them. - Assuming fairness between equal tasks. Two tasks that both yield every 2 ms still share the loop 50/50; if one is latency-sensitive, give the bulk work fewer, longer gaps or move it off the loop.
- Using
sleep(0.0001)"to be safe". Any non-zero delay creates a timer and is subject to clock resolution; in a tight loop it can cost orders of magnitude more thansleep(0). - Eager tasks. With the eager task factory, a new task runs synchronously until its first await; a CPU-heavy coroutine started eagerly stalls its creator, not just the loop.
Frequently Asked Questions¶
What does await asyncio.sleep(0) do?
It suspends the current task and reschedules it immediately, at the back of the event loop's ready queue. No timer is created. Every callback that became ready in the meantime — I/O completions, expired timers, other tasks — runs before the task resumes.
How often should I call sleep(0) in a CPU loop?
Yield on elapsed time rather than an iteration count: every 1–5 ms is a reasonable starting budget. In testing, yielding every 100 iterations of 3 µs work kept a neighbour's worst lag at 1.2 ms for a 4% throughput cost, while yielding every iteration cost 47%.
Is asyncio.sleep(0) expensive?
About 1.6 µs per call on Python 3.14 on Linux — cheap in absolute terms, but comparable to a small unit of Python work, so yielding on every iteration of a tight loop adds a large relative overhead.
Does sleep(0) make CPU-bound code faster?
No. It makes the event loop fairer by bounding how long other tasks wait, but the CPU work still runs on the loop thread and still takes the same total time. Large or long-running CPU work belongs in a process pool.
Can a task be cancelled while it is in a CPU loop?
Only at an await. A loop without awaits ignores cancel() and timeouts until it finishes; a periodic sleep(0) gives cancellation a place to land.
Related¶
- Task Scheduling & Lifecycle — up to the topic overview of how tasks are scheduled.
- Finding blocking calls with asyncio debug mode — locating the loops that need yields in the first place.
- Asyncio Fundamentals & Event Loop Architecture — the section overview.