Skip to content

Yielding Control with asyncio.sleep(0)

A coroutine that never awaits never lets anything else run. That is the whole contract of cooperative scheduling, and it bites hardest in code that looks async — an async def that parses a large payload, walks a big dictionary or scores a few thousand items in a loop. In a test, a loop doing 40,000 small units of work took 128 ms, and a heartbeat task that wanted to wake every 5 ms was late by 128 ms: the full duration of the loop. Adding await asyncio.sleep(0) every 100 iterations cut the heartbeat's worst lateness to 1.2 ms and cost about 4% in throughput. This guide covers what that one line does inside the event loop, how to choose the yield interval, and the point at which yielding stops being the right answer.

Prerequisites

  • Python 3.11+; every example is stdlib only. Timings below were taken on Python 3.14 on Linux.
  • The ready queue model from Task Scheduling & Lifecycle: a task runs until it awaits something that is not ready, then the loop picks the next callback.
  • A way to see lag — the heartbeat technique from measuring event loop lag in production is used throughout.

1. Understand what sleep(0) does

asyncio.sleep(0) is special-cased. It does not create a timer: it yields a bare None up through the task's __step, which tells the task machinery "reschedule me immediately". The task's next step goes to the back of the ready queue, behind every callback that became ready while it was running — I/O completions, timer expirations, other tasks' wake-ups.

import asyncio


async def chatty(name: str) -> None:
    for i in range(3):
        print(name, i)
        await asyncio.sleep(0)          # back of the ready queue, no timer created


async def main() -> None:
    async with asyncio.TaskGroup() as tg:
        tg.create_task(chatty("a"))
        tg.create_task(chatty("b"))


asyncio.run(main())
# a 0, b 0, a 1, b 1, a 2, b 2  — strict round-robin

Without the sleep(0), the output is a 0, a 1, a 2, b 0, b 1, b 2: task a runs to completion before b gets a turn. With it, the two tasks interleave one step at a time, because each yield puts the yielding task behind the other.

Any non-zero delay goes through a different path: a TimerHandle on the heap, a selector timeout, and the minimum resolution of the clock. sleep(0.001) in a hot loop therefore costs far more than sleep(0) and adds real latency to the task doing the work.

Verify: run the snippet and confirm strict alternation; remove the await and confirm one task runs to completion first.

One yield, one trip through the ready queue A sequence of 6 messages between 3 participants. One yield, one trip through the ready queue CPU task event loop heartbeat runs 100 work units await sleep(0): reschedule me poll selector with zero timeout timer due: wake heartbeat records lag, awaits sleep(0.005) next step of the CPU task The yield costs one pass through the loop; everything that became ready in between gets to run first.

2. Measure the stall before you fix it

Do not guess whether a loop needs yields. Run a heartbeat next to it and read the worst-case lateness:

import asyncio
import time


async def heartbeat(stop: asyncio.Event, lags: list[float]) -> None:
    while not stop.is_set():
        start = time.perf_counter()
        await asyncio.sleep(0.005)
        lags.append(time.perf_counter() - start - 0.005)


def work(i: int) -> int:                 # stand-in for parsing / scoring / transforming
    return sum(k * i for k in range(200))


async def run(yield_every: int, n: int = 40_000) -> tuple[float, float]:
    stop, lags = asyncio.Event(), []
    hb = asyncio.create_task(heartbeat(stop, lags))
    await asyncio.sleep(0.01)
    t0 = time.perf_counter()
    for i in range(n):
        work(i)
        if yield_every and i % yield_every == 0:
            await asyncio.sleep(0)
    elapsed = time.perf_counter() - t0
    stop.set()
    await hb
    return elapsed * 1000, max(lags) * 1000

The numbers from this harness, 40,000 iterations of about 3 µs each:

Yield every Loop total Worst heartbeat lag
never 127.5 ms 127.6 ms
1 iteration 187.4 ms 0.9 ms
10 iterations 143.8 ms 0.9 ms
100 iterations 132.7 ms 1.2 ms
1,000 iterations 135.2 ms 11.5 ms

Never yielding is the fastest way to finish the loop and the worst thing for everything else on the loop. Yielding on every iteration fixed the lag but added 47% to the loop's own runtime — at 1.6 µs per bare sleep(0), a yield costs about half of what one unit of work does.

Verify: your own numbers will differ; what matters is the shape — lag tracks the time between yields, and throughput loss tracks the number of yields.

Worst heartbeat lateness by yield interval 4 horizontal bars comparing never yields with the others. Worst heartbeat lateness by yield interval never yields 127.6 ms every 1,000 iterations 11.5 ms every 100 iterations 1.2 ms every 10 iterations 0.9 ms Python 3.14, Linux; each iteration is about 3 µs of pure-Python arithmetic. The lag a neighbour sees is the time between yields, so pick the interval from a latency budget.

3. Pick the interval from a time budget, not a count

An iteration count is only a proxy. The real knob is time between yields, and the right value is a fraction of your latency budget — if the service promises a p99 of 50 ms, a 10 ms stall is already a fifth of it. Yield on elapsed time and the loop stays correct when the work per item changes:

import asyncio
import time


async def cooperative(items, fn, budget_s: float = 0.002):
    """Apply fn to every item, yielding whenever budget_s of CPU time has passed."""
    out = []
    deadline = time.perf_counter() + budget_s
    for item in items:
        out.append(fn(item))
        if time.perf_counter() >= deadline:
            await asyncio.sleep(0)
            deadline = time.perf_counter() + budget_s
    return out

time.perf_counter() costs well under 100 ns, so checking it every iteration is cheap compared with a yield. A 2 ms budget bounds the stall any neighbour sees to roughly 2 ms plus one item, regardless of whether an item takes 3 µs or 300 µs.

Verify: run the heartbeat harness against cooperative() with items of very different cost; the worst lag should stay near the budget in every case.

4. Know when yielding is the wrong tool

Yielding fixes fairness, not cost. The CPU work still happens on the event loop thread, so the loop's total capacity for everything else is reduced by exactly the CPU time the work consumes. Three signs it is time to move the work off the loop:

  • The work is large relative to the request rate. A 128 ms computation per request at 20 requests per second is 2.5 s of CPU per second — no amount of yielding fits that on one thread.
  • The work calls into C that holds the GIL in one long call — a big json.loads, a regex over megabytes, hashlib on a large buffer. There is no Python loop to put a yield in.
  • Latency of the work itself matters. Yielding makes the work take longer, by design.
import asyncio
from concurrent.futures import ProcessPoolExecutor


async def score_batch(pool: ProcessPoolExecutor, batch: list[int]) -> list[int]:
    loop = asyncio.get_running_loop()
    return await loop.run_in_executor(pool, _score_all, batch)   # zero loop stall


def _score_all(batch: list[int]) -> list[int]:
    return [sum(k * i for k in range(200)) for i in batch]

The decision between threads and processes for that work is covered in offloading CPU work with run_in_executor; for pure-Python CPU work on a GIL build, it is a process pool.

Verify: with the work in a process pool, the heartbeat's worst lag stays at its idle level no matter how long the batch takes.

Yield on the loop, or move the work off it? A decision on How much CPU does it need per second with 3 outcomes. Yield on the loop, or move the work off it? How much CPU does it need per second? a few ms, in small steps yield on a time budget sleep(0) every ~2 ms one long C call to_thread or a process nothing to yield inside a large share of a core process pool the loop cannot fit it Yielding spreads CPU work fairly; it never makes the work cheaper.

5. Put the yield where cancellation can reach it

A task can only be cancelled at an await. A CPU loop with no awaits is uncancellable for its whole duration — task.cancel() sets a flag that is not acted on until the loop finishes. Each sleep(0) is also a cancellation point, which is a second reason to have them:

async def cancellable_transform(rows):
    done = 0
    try:
        for i, row in enumerate(rows):
            transform(row)
            done += 1
            if i % 100 == 0:
                await asyncio.sleep(0)       # CancelledError is raised here
    except asyncio.CancelledError:
        log.info("transform cancelled after %d rows", done)
        raise                                 # always re-raise

With a yield every 100 rows, a cancellation landed 0.03 ms after cancel() in the harness above instead of after the whole loop. The same reasoning applies to asyncio.timeout(): a timeout can only fire at an await, so a CPU loop without yields overruns its deadline by however long it runs. Implementing cooperative cancellation in CPU loops covers the pattern for work that has moved into threads, where sleep(0) is not available.

Verify: cancel the task 10 ms into the loop and confirm it finishes within a millisecond or two, with the partial count logged.

Verification

The change is working when:

  • A heartbeat beside the loop reports worst lag close to your yield budget instead of the loop's full duration.
  • Loop throughput drops by single-digit percent, not tens of percent — if it drops more, you are yielding too often.
  • Cancellation and timeouts take effect within one yield interval.
  • In debug mode (PYTHONASYNCIODEBUG=1), the "Executing took N seconds" warnings for this task disappear, since no single step exceeds slow_callback_duration any more.

Diagnostic Hook: export the heartbeat's worst lag per minute as a gauge and alert when it exceeds a quarter of your latency budget. Pair it with the slow-callback log from tracing slow callbacks: the log names the task, the gauge tells you how bad it is.

Pitfalls & edge cases

  • sleep(0) inside a lock. Yielding while holding an asyncio.Lock lets other tasks run but not the ones waiting for that lock; if they are the ones you wanted to unblock, nothing improved.
  • Yielding from a tight loop over a generator that itself blocks. If each next() does a blocking read, the yield happens between stalls rather than preventing them.
  • Assuming fairness between equal tasks. Two tasks that both yield every 2 ms still share the loop 50/50; if one is latency-sensitive, give the bulk work fewer, longer gaps or move it off the loop.
  • Using sleep(0.0001) "to be safe". Any non-zero delay creates a timer and is subject to clock resolution; in a tight loop it can cost orders of magnitude more than sleep(0).
  • Eager tasks. With the eager task factory, a new task runs synchronously until its first await; a CPU-heavy coroutine started eagerly stalls its creator, not just the loop.

Frequently Asked Questions

What does await asyncio.sleep(0) do?

It suspends the current task and reschedules it immediately, at the back of the event loop's ready queue. No timer is created. Every callback that became ready in the meantime — I/O completions, expired timers, other tasks — runs before the task resumes.

How often should I call sleep(0) in a CPU loop?

Yield on elapsed time rather than an iteration count: every 1–5 ms is a reasonable starting budget. In testing, yielding every 100 iterations of 3 µs work kept a neighbour's worst lag at 1.2 ms for a 4% throughput cost, while yielding every iteration cost 47%.

Is asyncio.sleep(0) expensive?

About 1.6 µs per call on Python 3.14 on Linux — cheap in absolute terms, but comparable to a small unit of Python work, so yielding on every iteration of a tight loop adds a large relative overhead.

Does sleep(0) make CPU-bound code faster?

No. It makes the event loop fairer by bounding how long other tasks wait, but the CPU work still runs on the loop thread and still takes the same total time. Large or long-running CPU work belongs in a process pool.

Can a task be cancelled while it is in a CPU loop?

Only at an await. A loop without awaits ignores cancel() and timeouts until it finishes; a periodic sleep(0) gives cancellation a place to land.