Skip to content

Comparing Context Switch Costs of Threads and Tasks

"Async is faster because switching tasks is cheaper than switching threads" is true, but the size of the difference is often overstated, and the factor that dominates in real services is a different one. Measured on Python 3.14 on a 24-core Linux machine: passing a value to another thread and back through queue.SimpleQueue took 7.2 µs per round trip and two operating-system context switches; the same exchange between two asyncio tasks through asyncio.Queue took 4.3 µs and no OS switches, and with bare futures 3.4 µs. Passing a token around a ring of 1,000 workers cost 7.65 µs per hop for threads and 3.07 µs for tasks. Those are differences of two to three times. The large difference appeared when one other thread was doing pure-Python CPU work: the thread round trip went from 5.9 µs to 5,096 µs at the median — a waiting thread needs the GIL, and the default switch interval is 5 ms. asyncio was hit too: with the same busy thread in the process, the task round trip p99 rose from 6.9 µs to 5,146 µs. This guide measures each of these and shows what to change.

Prerequisites

1. Measure a two-party handoff

The smallest unit of switching is a ping-pong: one side sends a value, the other sends it back. With threads:

def threads_round_trip(n=100_000):
    a, b = queue.SimpleQueue(), queue.SimpleQueue()

    def pong():
        for _ in range(n):
            b.put(a.get())

    t = threading.Thread(target=pong)
    t.start()
    start = time.perf_counter()
    for i in range(n):
        a.put(i)
        b.get()
    elapsed = time.perf_counter() - start
    t.join()
    return elapsed / n

And with tasks:

async def tasks_round_trip(n=100_000):
    a, b = asyncio.Queue(), asyncio.Queue()

    async def pong():
        for _ in range(n):
            b.put_nowait(await a.get())

    t = asyncio.create_task(pong())
    start = time.perf_counter()
    for i in range(n):
        a.put_nowait(i)
        await b.get()
    elapsed = time.perf_counter() - start
    await t
    return elapsed / n

Measured over three runs each: threads with SimpleQueue, 7.15–7.63 µs per round trip; threads with a pair of threading.Event objects, 11.0–11.3 µs; tasks with asyncio.Queue, 4.25–4.49 µs; tasks resolving futures directly, 3.38–3.52 µs. getrusage showed why threads cost more: 2.0 operating-system context switches per round trip, against effectively zero for tasks — between 1 and 18 in total across 100,000 round trips. A task switch is a return to the event loop and a call into another coroutine, all inside one OS thread.

Verify: count ru_nvcsw + ru_nivcsw from resource.getrusage around the loop; thread handoffs show about one OS switch per direction, task handoffs show none.

Round trip between two workers, Python 3.14 4 horizontal bars comparing threads, threading.Event with the others. Round trip between two workers, Python 3.14 threads, threading.Event 11.3 us threads, queue.SimpleQueue 7.2 us tasks, asyncio.Queue 4.3 us tasks, bare futures 3.4 us Threads: 2.0 OS context switches per round trip. Tasks: none. Median of three runs of 100,000 round trips.

2. Scale the number of workers

Real services have more than two workers. A ring — each worker receives a token and passes it to the next — shows how cost per switch changes with worker count:

async def node(i, qs, hops):
    while True:
        v = await qs[i].get()
        if v < 0:
            qs[(i + 1) % len(qs)].put_nowait(v)
            return
        qs[(i + 1) % len(qs)].put_nowait(v + 1 if v + 1 < hops else -1)

The thread version is identical with queue.SimpleQueue and threading.Thread. Measured over 200,000 hops: with 2 workers, threads 3.98 µs per hop and tasks 2.37 µs; with 100, threads 4.54 µs and tasks 2.22 µs; with 1,000, threads 7.65 µs and tasks 3.07 µs. Thread cost rose by 92% from 2 to 1,000 workers as the kernel scheduled across more runnable threads; task cost rose by 30%. Pinning the two-thread case to a single core with taskset made it worse — 6.45 µs per hop and 3.0 OS switches per hop, as threads competed for the same CPU — while the task ring was unchanged at 2.65 µs. On the free-threaded 3.14 build, without a GIL, the thread ring ran at 4.64 µs per hop for 2 workers and 5.22 µs for 100: removing the GIL did not make handoffs cheaper.

Verify: run the ring at your real worker count; the ratio, not the absolute numbers, is what transfers between machines.

3. Find the cost of a bare yield

For tasks, the cheapest switch is await asyncio.sleep(0), which yields to the loop and resumes on the next iteration. Measured with K tasks each yielding in a loop for one second:

async def spin():
    while not stop:
        await asyncio.sleep(0)
        count[0] += 1

One task: 1.61 µs per yield, since every yield also runs a full loop iteration including the selector poll. With 10 tasks, 0.76 µs; with 100 and 1,000, 0.67 µs, as one loop iteration processed many ready tasks; with 10,000, 1.78 µs, as the ready queue and task objects stopped fitting in CPU caches. A coroutine that yields every 0.1 ms to stay responsive — the advice in yielding control with asyncio.sleep(0) — spends under 2% of its time switching.

Verify: measure yields per second at your real task count; above about 10,000 runnable tasks, expect per-switch cost to rise.

Per-hop cost in a ring of workers A grid of 4 rows by 4 columns. Per-hop cost in a ring of workers workers threads, GIL build tasks threads, free-threaded 2 3.98 us 2.37 us 4.64 us 100 4.54 us 2.22 us 5.22 us 1,000 7.65 us 3.07 us not measured 2, one core 6.45 us, 3.0 OS switches/hop 2.65 us not measured 200,000 hops; Python 3.14.4 GIL build and 3.14.6 free-threaded.

4. Measure what one busy thread does to switching

The numbers above assume nothing else wants the CPU. In a real process, something often does — a JSON encode, a template render, a hash, running in a thread pool. With one thread spinning in pure Python:

def burn():
    x = 0
    while not stop:
        x += 1

threading.Thread(target=burn).start()
# ... then the same two-thread ping-pong, timing each round trip

Measured with 2,000 round trips: the median went from 5.9 µs to 5,096 µs, and the p99 from 16.6 µs to 15,235 µs. A thread woken by SimpleQueue still has to take the GIL, and the busy thread only releases it when the interpreter's switch interval expires — 5 ms by default, sys.getswitchinterval(). Every handoff waited one interval. Lowering it with sys.setswitchinterval(0.0005) cut the median to 572 µs and the p99 to 2,187 µs. On the free-threaded build the busy thread made no difference: 6.7 µs median, 16.2 µs p99.

asyncio is not immune. With the same busy thread in the process, the task ping-pong — which never leaves the event loop thread — went from a 4.0 µs median and 6.9 µs p99 to an 8.2 µs median and 5,146 µs p99, and the total for 20,000 round trips rose from 0.09 s to 12.6 s. Each time the loop thread lost the GIL, it waited a full interval to get it back.

Verify: run the ping-pong with and without a CPU-bound thread; a p99 near sys.getswitchinterval() means GIL contention dominates switching cost.

p99 round trip with one CPU-bound thread in the process 4 horizontal bars comparing threads, switch interval 5 ms with the others. p99 round trip with one CPU-bound thread in the process threads, switch interval 5 ms 15,235 us asyncio tasks, switch interval 5 ms 5,146 us threads, switch interval 0.5 ms 2,187 us threads, free-threaded build 16.2 us Idle process: threads 16.6 us p99, tasks 6.9 us p99. GIL waits, not context switches, set the latency here.

5. Apply the numbers to a design choice

For an I/O-bound service, switching cost is rarely the deciding factor. At 7 µs per thread round trip, 10,000 handoffs per second cost 70 ms of CPU per second — 7% of a core; tasks at 4 µs cost 4%. That difference matters at very high message rates, where it adds up alongside the other per-request costs compared in when asyncio is slower than threads. What matters more is where CPU work runs:

# Keep pure-Python CPU work out of the process that handles latency-sensitive I/O
pool = concurrent.futures.ProcessPoolExecutor(max_workers=4)

async def handler(payload):
    return await asyncio.get_running_loop().run_in_executor(pool, render_report, payload)

A CPU-bound function in asyncio.to_thread keeps the event loop technically free, but, as step 4 measured, it still holds the GIL for 5 ms at a time, and every coroutine on the loop waits for it. A process pool moves the GIL contention out of the serving process. On a free-threaded build, threads stop competing for a GIL, and the trade-off changes; see evaluating free-threaded Python for CPU-bound threads.

Verify: latency percentiles for I/O requests stay flat while CPU-heavy requests run; if p99 tracks the switch interval, move the CPU work to processes.

Verification

The comparison is useful when:

  • Handoff cost is measured on your Python and hardware, not assumed — here 7.2 µs for threads and 4.3 µs for tasks.
  • OS context switches are counted, confirming that task switches stay inside one thread.
  • The test is repeated with a CPU-bound thread present, because that changes the result by three orders of magnitude.
  • Pure-Python CPU work runs in processes in latency-sensitive services on GIL builds.

Diagnostic Hook: when tail latency in a threaded or asyncio service clusters at multiples of 5 ms, compare it with sys.getswitchinterval(). That pattern is GIL contention from a CPU-bound thread, not slow switching or slow I/O.

Pitfalls & edge cases

  • Choosing asyncio for switch cost alone. Measured: about 2× cheaper, not 100×.
  • CPU work in to_thread. Measured: task p99 rose to 5.1 ms.
  • Pinning threads to one core. Measured: 6.45 µs per hop against 3.98 µs.
  • Lowering the switch interval as a fix. It reduced p99 to 2.2 ms but does not remove the wait.

Frequently Asked Questions

Is an asyncio task switch faster than a thread switch?

Yes, about two times in these measurements: 4.3 µs per round trip through asyncio.Queue against 7.2 µs through queue.SimpleQueue, because tasks switch without the OS scheduler.

How many OS context switches does a thread handoff cost?

Measured with getrusage: 2.0 per round trip, one per direction, and 3.0 per hop when both threads were pinned to one core. Task handoffs caused effectively none.

Why does my thread handoff take 5 ms?

Another thread is running pure-Python CPU work and holds the GIL until the 5 ms switch interval expires. One busy thread raised the median round trip from 5.9 µs to 5,096 µs.

Does free-threaded Python make thread switches cheaper?

No: a 2-thread ring took 4.64 µs per hop against 3.98 µs with the GIL. It removes GIL waits instead — with a busy thread, p99 stayed at 16.2 µs.