Skip to content

Cancelling asyncio.to_thread Calls

asyncio.to_thread makes blocking code awaitable, and the await is cancellable like any other: cancel the task and CancelledError arrives at once. The thread, though, cannot be interrupted — Python has no safe way to stop a running thread — so the blocking function carries on to the end. Tested on Python 3.14: a task awaiting a 2-second blocking call was cancelled after 0.1 s and the await raised CancelledError after 0.10 s, but the function ran to completion at 2.0 s. With a two-thread executor and two such cancelled calls still running, the next to_thread call waited 1.90 s for a free thread. And a script whose main() returned at 0.10 s after cancelling a 3-second thread call did not exit until 3.00 s, because asyncio.run waits for the default executor's threads at shutdown. Passing a threading.Event that the function checks stopped a cooperative version after 0.15 s. This guide makes thread calls stop when their caller does.

Prerequisites

1. Know what cancellation does and does not stop

Cancelling the awaiting task resolves the await immediately; the function in the thread is unaffected:

def blocking_job() -> None:
    time.sleep(2.0)                        # or a long computation, or a blocking SDK call
    write_result()                          # still happens after the caller was cancelled


task = asyncio.create_task(asyncio.to_thread(blocking_job))
await asyncio.sleep(0.1)
task.cancel()
await asyncio.gather(task, return_exceptions=True)   # measured: returns after 0.10 s
# ...and blocking_job still finished at 2.0 s, including write_result()

That has three consequences. Side effects after the cancellation point still happen — a write the caller thought it abandoned completes anyway. The thread stays busy, occupying a slot in the executor. And CPU or I/O keeps being spent on work nobody is waiting for. For short calls these are harmless; for long ones they decide whether cancellation means anything.

Verify: add a log line at the end of the threaded function; it still appears after the caller's cancellation.

What a cancelled to_thread call costs 4 horizontal bars comparing await raises CancelledError with the others. What a cancelled to_thread call costs await raises CancelledError 0.10 s uncooperative function finishes 2.00 s next call waits for a thread (pool of 2) 1.90 s cooperative function stops 0.15 s Python 3.14; blocking function sleeps 2 s in 50 ms steps; cancel sent at 0.1 s. Cancelling the await is instant; stopping the work needs the function's help.

2. Make the function cooperative with a stop event

The thread can stop if the function checks a flag between steps. Pass a threading.Event and set it when the awaiting task is cancelled:

import threading


def process_rows(rows, stop: threading.Event) -> int:
    done = 0
    for row in rows:
        if stop.is_set():
            return done                     # stop between units of work
        handle(row)
        done += 1
    return done


async def run_cancellable(func, *args) -> object:
    stop = threading.Event()
    try:
        return await asyncio.to_thread(func, *args, stop)
    except asyncio.CancelledError:
        stop.set()                          # tell the thread to finish early
        raise

Measured: with the event, the function noticed the cancellation and returned after 0.15 s instead of running for 2.0 s. The granularity is the step between checks — every row, every chunk, every page — so choose steps short enough for the responsiveness you need. threading.Event is safe to set from the event loop thread and read from the worker. The same idea for CPU loops inside asyncio itself is in implementing cooperative cancellation in CPU loops.

Verify: cancelling the caller makes the function return within one step, shown by its log line.

3. Bound uncooperative calls with library timeouts

Code you do not control — a blocking SDK call, a database driver, a subprocess wait — cannot check your event. Use its own timeout, which actually stops the blocking operation:

# Network client: timeouts end the blocking call inside the thread
response = await asyncio.to_thread(requests.get, url, timeout=(3, 10))

# boto3: connect/read timeouts in the client config
s3 = boto3.client("s3", config=Config(connect_timeout=3, read_timeout=10))
data = await asyncio.to_thread(lambda: s3.get_object(Bucket=b, Key=k)["Body"].read())

# Locks and queues in threads: always with a timeout
item = await asyncio.to_thread(work_queue.get, True, 5.0)

An asyncio.timeout around the await bounds how long the caller waits; the library timeout bounds how long the thread is occupied. You want both, with the library timeout shorter or equal, so a cancelled call frees its thread soon after. Calls with no timeout at all — socket.recv without a timeout, queue.get() without one — can occupy a thread forever.

Verify: every blocking call run through to_thread has an explicit timeout argument or config.

4. Keep cancelled calls from starving the executor

Threads still running cancelled work count against the executor's size. Tested with a two-thread executor: after two long calls were cancelled, a new short call waited 1.90 s for one of them to finish. In a service, a burst of cancellations — clients disconnecting, timeouts firing — can fill the default executor and stall every other to_thread user:

from concurrent.futures import ThreadPoolExecutor

reports_executor = ThreadPoolExecutor(max_workers=8, thread_name_prefix="reports")


async def build_report(params) -> bytes:
    loop = asyncio.get_running_loop()
    stop = threading.Event()
    try:
        return await loop.run_in_executor(reports_executor, render_report, params, stop)
    except asyncio.CancelledError:
        stop.set()
        raise

A dedicated executor for long or risky work isolates it: cancelled reports can only occupy report threads, and the default executor stays free for short calls such as file reads and DNS lookups. Export the number of busy threads per executor; a pool that stays full while callers are cancelling means threads are running abandoned work.

Verify: under a burst of cancelled long calls, short to_thread calls elsewhere keep their normal latency.

A to_thread call that stops when its caller does A flow of 5 stages. A to_thread call that stops when its caller does dedicated executor isolates long work pass stop Event checked between steps caller cancelled set the event, re-raise function returns early thread freed library timeouts bound uncheckable steps The await is cancelled by asyncio; the thread is stopped by your code.

5. Account for shutdown waiting on threads

asyncio.run shuts down the default executor before returning, and that waits for running threads. Tested: a script whose main() returned at 0.10 s after cancelling a 3-second thread call exited only at 3.00 s. For services, a long uncooperative call can stretch shutdown past the orchestrator's grace period:

async def main() -> None:
    stop_all = threading.Event()
    try:
        await serve(stop_all)
    finally:
        stop_all.set()                                  # cooperative functions stop now
        await asyncio.to_thread(reports_executor.shutdown, wait=True, cancel_futures=True)

Signalling every cooperative function through a shared stop event, and shutting down dedicated executors with cancel_futures=True (which drops queued work that has not started), keeps shutdown short. Uncooperative calls still finish their current step, so their library timeouts set the worst case. Threads cannot be killed; if a call can hang indefinitely, run it in a subprocess, which can be, as in killing subprocess trees on cancellation.

Verify: process shutdown time stays below the grace period even with long thread calls in progress.

How should this blocking call be made cancellable? A decision on What runs in the thread with 4 outcomes. How should this blocking call be made cancellable? What runs in the thread? your own loop stop Event checked per step stops in one step library call library timeout thread freed on timeout long or bursty work dedicated executor protects default pool may hang forever subprocess instead can be killed Threads end only when their code returns; plan how it will.

Verification

Thread calls stop when their callers do when:

  • Cooperative functions check a stop event set on cancellation.
  • Library calls in threads have explicit timeouts.
  • Long or bursty work uses a dedicated executor.
  • Shutdown signals stop events and bounds executor shutdown.

Diagnostic Hook: export busy threads per executor and count cancellations of to_thread awaits. Busy threads that stay high after a burst of cancellations mean abandoned work is still running; a slow process exit with an idle event loop means asyncio.run is waiting for executor threads.

Pitfalls & edge cases

  • Assuming cancel stops the thread. Measured: the function ran to completion.
  • Side effects after the cancel point. They still happen unless the function checks a flag.
  • Cancelled calls filling the pool. Measured: a new call waited 1.90 s.
  • Shutdown waiting on threads. Measured: exit at 3.00 s instead of 0.10 s.

Frequently Asked Questions

Does cancelling asyncio.to_thread stop the thread?

No. The await raises CancelledError immediately, but the function keeps running to completion; in testing a cancelled 2-second call still finished at 2.0 s.

How do I stop a function running in asyncio.to_thread?

Pass it a threading.Event, check the event between units of work, and set it in an except CancelledError block before re-raising. In testing that stopped the function after 0.15 s.

Why does my asyncio program take long to exit after cancelling tasks?

asyncio.run waits for the default executor's threads at shutdown. A cancelled 3-second thread call delayed exit until 3.0 s in testing.

Can I kill a thread in Python?

Not safely. Make the code cooperative, give library calls timeouts, or run work that may hang in a subprocess, which can be killed.