Skip to content

Cancelling Streaming LLM Responses

Users abandon answers all the time — they close the tab, press stop, or ask a new question before the old one finishes. Code abandons them too: an early exit once enough text has arrived, a timeout, a task cancelled during shutdown. What matters is whether the model server stops generating, because output tokens cost compute — and, depending on the provider's billing, money — whether or not anyone reads them. The server only learns that the client has gone when the HTTP connection carrying the stream is closed. To measure which client patterns close it, the anthropic SDK (1.11.0) on Python 3.14 requested 2,000-token streams from a local mock of the Messages API producing a token every 5 ms, and the client stopped after about 100 tokens. With messages.stream() used as a context manager, the server had sent 100 tokens when it saw the disconnect. Cancelling the consuming task stopped it at 108. Breaking out of an iterator from messages.create(stream=True) and calling await stream.close() stopped it at 98. Breaking out of the same iterator without closing it left the server streaming — 491 tokens two seconds later, and all 2,000 after twelve. This guide makes every exit path close the stream.

Prerequisites

1. Consume streams inside a context manager

The SDK's streaming helper is an async context manager; leaving it closes the underlying HTTP response, whatever the reason for leaving:

async def first_paragraph(prompt: str) -> str:
    out = []
    async with client.messages.stream(model="claude-sonnet-5-5", max_tokens=2000,
                                      messages=[{"role": "user", "content": prompt}]) as stream:
        async for text in stream.text_stream:
            out.append(text)
            if "\n\n" in text:
                break                            # leaving the block closes the HTTP stream
    return "".join(out)

Measured: the client stopped after 0.56 s and the mock recorded a disconnect after exactly 100 tokens; two seconds later its count had not moved. Exceptions raised inside the block, return, break and cancellation all run the context manager's exit, so the same code is correct for every way of stopping early. This is the single most important habit for LLM streaming in async code, because the failure it prevents — paying for and computing output nobody reads — is invisible in your own logs.

Verify: after an early exit, the provider's usage for that request (or the mock's counter) shows output tokens close to what you read, not max_tokens.

Tokens the server generated when the client stopped at ~100 4 horizontal bars comparing messages.stream() context, break with the others. Tokens the server generated when the client stopped at ~100 messages.stream() context, break 100 task cancelled mid-stream 108 create(stream=True), break, await close() 98 create(stream=True), break, never closed all 2,000 Local mock of the Messages API at 5 ms per token; anthropic 1.11.0, Python 3.14. The unclosed stream kept generating long after the client stopped reading.

2. Close raw streams explicitly

messages.create(stream=True) returns an AsyncStream that is not tied to a block. Iterating it to the end closes it; leaving early does not:

# Leaks: the response stays open after break
stream = await client.messages.create(model=MODEL, max_tokens=2000, messages=msgs, stream=True)
async for event in stream:
    if done(event):
        break


# Correct: close in finally, or use the stream as a context manager
stream = await client.messages.create(model=MODEL, max_tokens=2000, messages=msgs, stream=True)
try:
    async for event in stream:
        if done(event):
            break
finally:
    await stream.close()


async with await client.messages.create(model=MODEL, max_tokens=2000, messages=msgs, stream=True) as stream:
    async for event in stream:
        if done(event):
            break

Measured: the leaked stream showed no disconnect, kept the request in flight, and the server produced all 2,000 tokens over the following 12 seconds — 20 times what the client read. Holding a reference to the stream, as code that stores it on an object or in a list does, prevents garbage collection from closing it either. With await stream.close() the server stopped at 98 tokens.

Verify: a lint rule or code review checks that every create(..., stream=True) is used with async with or closed in finally.

3. Rely on cancellation reaching the stream

When the task that consumes a stream is cancelled — a request timeout, a client disconnect in a web handler, application shutdown — the CancelledError is raised at the await inside the iteration, propagates out through the async with, and the exit closes the response:

async def consume(prompt: str) -> None:
    async with client.messages.stream(model=MODEL, max_tokens=2000,
                                      messages=[{"role": "user", "content": prompt}]) as stream:
        async for text in stream.text_stream:
            await websocket.send_text(text)


task = asyncio.create_task(consume("long essay"))
await asyncio.sleep(0.6)
task.cancel()                                  # server saw the disconnect at 108 tokens

Measured: cancelling 0.6 s into the stream stopped the server at 108 tokens. Cancellation only helps if it is not swallowed on the way out: an except Exception around the iteration does not catch CancelledError, but an except BaseException or a bare except: that does not re-raise would leave the task running — see preventing CancelledError leaks in cleanup. The same mechanism makes asyncio.timeout around the consuming code stop generation at the deadline.

Verify: cancelling a consumer task in a test produces a disconnect at the server within one token interval.

How a cancellation reaches the model server A sequence of 5 messages between 4 participants. How a cancellation reaches the model server caller consumer task SDK stream model server task.cancel() CancelledError at await in text_stream __aexit__: close the HTTP response TCP connection closed generation stops (108 tokens sent) Every step depends on the stream being owned by a context manager in the cancelled task.

4. Do not detach the stream from its consumer

A common structure for fan-out or relaying puts the stream in one task and the consumer in another, connected by a queue. When the consumer goes away, the producer does not know:

# Leaks: the producer outlives its consumer
async def producer(q: asyncio.Queue) -> None:
    async with client.messages.stream(**req) as stream:
        async for text in stream.text_stream:
            await q.put(text)                    # an unbounded queue never blocks
    await q.put(None)

asyncio.create_task(producer(q))                 # nobody cancels this when the consumer leaves

In the relay test this pattern kept the upstream stream alive after the browser disconnected: 441 tokens had been generated two seconds after the client left, and generation was still running. Two fixes: keep the stream in the consumer's own task (no queue), or tie the producer's lifetime to the consumer with a TaskGroup or an explicit cancel() in the consumer's finally. A bounded queue also helps — the producer blocks when the consumer stops reading — but blocking is not closing; the HTTP stream stays open until something cancels the producer. The relay version of this is in relaying LLM token streams through FastAPI.

Verify: for every background task that owns a stream, find the code that cancels it when its consumer goes away.

5. Stop at a token budget, not just on demand

Cancellation on demand covers users who leave; a token budget covers answers that run long. max_tokens is the hard limit; an application-level budget can stop earlier with the same close-on-exit mechanics:

async def bounded_answer(prompt: str, budget_tokens: int) -> str:
    out: list[str] = []
    async with client.messages.stream(model=MODEL, max_tokens=4096,
                                      messages=[{"role": "user", "content": prompt}]) as stream:
        async for event in stream:
            if event.type == "content_block_delta" and event.delta.type == "text_delta":
                out.append(event.delta.text)
                if len(out) >= budget_tokens:      # chunks approximate tokens
                    break
    return "".join(out)

Chunk counts approximate tokens; exact usage arrives with the message_delta event at the end of a completed stream, so a budget enforced mid-stream has to count chunks. Setting max_tokens to what a task genuinely needs is simpler than stopping it later. The early-exit check is for budgets that vary per request — a user's remaining quota, a time-boxed summary — which max_tokens cannot express after the request is sent.

Verify: a request with a small budget stops within a few chunks of it, and the server's reported output tokens match.

How does this stream get closed when we stop early? A decision on Who owns the stream with 4 outcomes. How does this stream get closed when we stop early? Who owns the stream? the consuming code, via messages.stream() context manager exit 100 tokens the consuming code, via create(stream=True) async with, or close() in finally 98 vs 2,000 a background producer task cancel it with its consumer 441 and counting otherwise nobody stops, answer runs long max_tokens or a budget break bounded by design The server stops when the connection closes, and only then.

Verification

Streams stop when their readers do when:

  • Every stream is consumed inside async with, or closed in finally.
  • Cancelling the consuming task closes the stream, with no handler swallowing CancelledError.
  • No background task owns a stream without being cancelled when its consumer leaves.
  • Early exits are visible in usage: output tokens close to what was read, not max_tokens.

Diagnostic Hook: compare, per request, the number of text chunks your code consumed with the output tokens the provider reports. Requests where reported tokens greatly exceed consumed chunks are streams that kept generating after you stopped reading — find the code path that exits without closing.

Pitfalls & edge cases

  • break out of create(stream=True) without close(). Measured: all 2,000 tokens generated after reading 100.
  • Producer tasks feeding queues. Measured: 441 tokens and still streaming after the reader left.
  • Swallowing CancelledError. The stream stays open and the task keeps running.
  • Large max_tokens as a default. It raises the cost of every stream that is not closed.

Frequently Asked Questions

How do I cancel a streaming Claude response in Python?

Leave the async with client.messages.stream(...) block — by break, return, exception or task cancellation — and the SDK closes the HTTP stream; in testing the server stopped within a few tokens. With create(stream=True), use async with or await stream.close().

Does the LLM keep generating if I stop reading the stream?

Yes, until the connection closes. An abandoned create(stream=True) iterator that was never closed let the server generate all 2,000 tokens after the client read 100.

Does task.cancel() stop an LLM stream?

Yes, if the stream is owned by a context manager in that task: cancelling 0.6 s in stopped the server at 108 tokens. A stream owned by a separate producer task keeps running unless that task is cancelled too.

What happens to tokens generated after the client disconnects?

That depends on the provider's billing, but generation itself only stops when the connection closes, so close streams promptly. Compare consumed chunks with reported usage to find streams that were not closed.