Skip to content

Streaming LLM Tokens Asynchronously in Python

Streaming is the default way to call a large language model from an interactive application: the first words arrive in a fraction of a second instead of after the whole answer is generated. From asyncio it is also the place where the usual HTTP assumptions break. A response can stay open for minutes, the first byte can take longer than a typical read timeout, errors can arrive after a 200 OK, and every token is a small event the event loop must parse. To measure these behaviours, the official anthropic Python SDK (1.11.0) and plain httpx were run on Python 3.14 against a local mock of the Messages API that streams server-sent events in the same format. With 50 concurrent streams of 400 tokens each, the client spent 52–54 µs of CPU per token parsing SSE with httpx and 77–97 µs through the SDK — about half a core to one core per 10,000 tokens per second. When the first token took 6 s, httpx's default timeout failed with ReadTimeout after 5.0 s while the SDK's 600 s read timeout let it finish in 6.5 s. A per-request timeout=3.0 that expired mid-stream raised httpx2.ReadTimeout, not the SDK's APITimeoutError. And an error event sent after streaming had begun surfaced as APIStatusError with status code 200. This guide streams tokens with those edges handled.

Prerequisites

1. Stream with the SDK's context manager

The SDK's messages.stream() helper opens the request, parses the server-sent events, and accumulates the final message for you:

import asyncio
from anthropic import AsyncAnthropic

client = AsyncAnthropic()                       # one client per process, reused


async def answer(prompt: str) -> str:
    parts: list[str] = []
    async with client.messages.stream(
        model="claude-sonnet-5-5",
        max_tokens=1024,
        messages=[{"role": "user", "content": prompt}],
    ) as stream:
        async for text in stream.text_stream:
            parts.append(text)                  # or forward it to the user
        final = await stream.get_final_message()
    log.info("tokens in=%d out=%d stop=%s", final.usage.input_tokens,
             final.usage.output_tokens, final.stop_reason)
    return "".join(parts)

Against the mock, with 200 ms to the first token and 10 ms per token after it, the first text arrived after 222 ms and 200 tokens in 2.24 s; the final message reported 200 output tokens and the stop reason. The context manager is not decoration: leaving it — normally or by exception — closes the HTTP response, which is what tells the server to stop generating. The cancellation guide shows what happens without it.

Verify: the usage numbers from get_final_message() are logged for every call; usage is what token budgets and costs are computed from.

The events of one streamed response A sequence of 7 messages between 2 participants. The events of one streamed response client API POST /v1/messages, stream: true 200 OK + message_start content_block_start (after TTFT) content_block_delta x N (one per chunk) content_block_stop message_delta: stop_reason, usage message_stop The status code is settled before the first token; failures after it arrive as events.

2. Budget CPU for the token rate

Every streamed chunk is an SSE event: bytes to decode, a line to split, JSON to parse, an object to build. On one event loop that adds up. Measured with 50 concurrent streams of 400 tokens at 5 ms per token — 20,000 tokens in about 2.3 s — two client implementations were compared:

# Raw httpx: parse SSE lines yourself
async with http.stream("POST", "/v1/messages", json=req | {"stream": True}, headers=headers) as r:
    r.raise_for_status()
    async for line in r.aiter_lines():
        if line.startswith("data: "):
            event = json.loads(line[6:])
            if event["type"] == "content_block_delta":
                handle(event["delta"]["text"])

The raw parser used 1.05–1.09 s of CPU (52–54 µs per token) with event-loop lag p99 of 2.4–3.3 ms. The SDK used 1.54–1.94 s (77–97 µs per token) with lag p99 of 4.2–12.2 ms across two runs — it builds typed event objects and an accumulated message snapshot. Neither is a problem for one user; at 10,000 tokens per second across all streams, the SDK alone would occupy most of a core. Services that relay many simultaneous streams should measure their own token rate, keep JSON work light per event, and run several worker processes, as in sizing uvicorn workers for async services.

Verify: at your peak concurrent streams, event-loop lag p99 stays below the latency your users notice.

Client CPU per streamed token, 50 concurrent streams 4 horizontal bars comparing raw httpx + json.loads, run 1 with the others. Client CPU per streamed token, 50 concurrent streams raw httpx + json.loads, run 1 52 us raw httpx + json.loads, run 2 54 us anthropic SDK stream, run 1 97 us anthropic SDK stream, run 2 77 us 20,000 tokens per run against a local mock streaming 200 tokens/s per stream; Python 3.14. Roughly half a core to a full core per 10,000 tokens per second, on the event loop.

3. Set timeouts for slow first tokens

A model may take seconds before the first token — long prompts, queued capacity, extended reasoning. A client timeout tuned for ordinary APIs kills those requests:

# httpx's default is 5 s for connect, read, write and pool
async with httpx.AsyncClient(base_url=URL) as http:
    async with http.stream("POST", "/v1/messages", json=req) as r:
        async for line in r.aiter_lines():       # ReadTimeout after 5.0 s if TTFT is 6 s
            ...

# The SDK's default: connect 5 s, read/write/pool 600 s
client = AsyncAnthropic()                        # completed the same request in 6.5 s

Measured against the mock with a 6-second time to first token: raw httpx with defaults raised ReadTimeout at 5.0 s; the SDK completed in 6.5 s. The SDK's read timeout is per read, not per request, so a stream that keeps producing tokens can run far longer than 600 s; to bound the whole call, put a deadline around it, which is the subject of the next step. Connect timeouts can stay short — 5 seconds is generous for establishing TLS to an API endpoint — and should, so an unreachable endpoint fails fast.

Verify: your read timeout exceeds the p99 time to first token observed in production, and a separate deadline bounds the total duration.

4. Catch the exceptions streaming actually raises

The SDK maps HTTP errors that happen before streaming starts to its own exception hierarchy. Failures after the response has started arrive differently, and two of them surprised in testing:

import httpx2                                   # the SDK's HTTP layer in anthropic 1.11.0
import anthropic


async def stream_text(prompt: str):
    """Yield text chunks; translate every failure into the application's own errors."""
    try:
        async with client.messages.stream(model=MODEL, max_tokens=1024,
                                          messages=[{"role": "user", "content": prompt}]) as s:
            async for text in s.text_stream:
                yield text
    except anthropic.APIStatusError as exc:                    # includes mid-stream error events
        raise LLMError(exc.status_code, str(exc)) from exc
    except (anthropic.APIConnectionError, httpx2.TransportError) as exc:
        raise LLMUnavailable(str(exc)) from exc


async def answer(prompt: str, deadline: float = 120.0) -> str:
    parts: list[str] = []
    async with asyncio.timeout(deadline):                      # the deadline lives in the caller
        async for text in stream_text(prompt):
            parts.append(text)
    return "".join(parts)

Measured: with the SDK's per-request timeout=3.0 and a 6-second gap mid-stream, the exception was httpx2.ReadTimeout, whose base classes are TimeoutException, TransportError, RequestError and HTTPError — it was neither anthropic.APITimeoutError nor anthropic.APIConnectionError, and an except anthropic.APIError would have let it escape. When the mock sent an error event with overloaded_error after 20 tokens, the SDK raised APIStatusError whose status_code was 200 — the status of the response that had already begun. asyncio.timeout(3.0) around the stream raised TimeoutError at exactly 3.0 s, independent of the SDK's internals. Where that timeout goes matters: placed inside an async generator around its yields, it misfired. With a consumer that awaited 50 ms per token, the deadline expired while the consumer — not the generator — was running, so the cancellation landed in the consumer's code and surfaced as a bare CancelledError after 5 tokens instead of a TimeoutError; with a fast consumer the same code raised TimeoutError as intended. Keep the deadline in the code that iterates, as above. The retry guide covers which of these to retry.

Verify: a test that injects a mid-stream timeout and a mid-stream error event gets your application's own exception types, not a raw transport error.

What each failure looked like to the caller A grid of 5 rows by 3 columns. What each failure looked like to the caller situation exception when 6 s TTFT, httpx defaults httpx ReadTimeout 5.0 s 6 s TTFT, SDK defaults none, completed 6.5 s SDK timeout=3.0, gap mid-stream httpx2.ReadTimeout (not an SDK error) 3.0 s asyncio.timeout(3.0) around the stream TimeoutError 3.0 s error event after 20 tokens APIStatusError, status_code 200 0.10 s Local mock of the Messages API; anthropic 1.11.0 on Python 3.14.

5. Reuse one client and keep streams independent

AsyncAnthropic owns an HTTP connection pool — up to 1,000 connections, 100 kept alive, by default — so create it once and share it across tasks. Each stream is an independent request on that pool, and concurrent streams do not block each other:

from contextlib import asynccontextmanager


@asynccontextmanager
async def lifespan(app):
    app.state.llm = AsyncAnthropic(max_retries=2)
    try:
        yield
    finally:
        await app.state.llm.close()             # closes pooled connections

Creating a client per request discards its pool and pays TLS setup on every call, the same cost measured for other HTTP clients in reusing aiohttp ClientSession across requests. The SDK's limits are generous; your concurrency limit should come from the provider's rate limits instead, as in fanning out LLM calls with bounded concurrency.

Verify: the number of client objects created over the life of the process is one per configuration, not one per request.

Verification

Async LLM streaming is production-ready when:

  • Streams are consumed inside messages.stream(), and usage is logged from the final message.
  • Read timeouts exceed the p99 time to first token, and asyncio.timeout bounds each call's total duration.
  • Mid-stream failures are caught, including httpx2.TransportError and APIStatusError with status 200.
  • CPU per token is known, and event-loop lag stays acceptable at peak concurrent streams.

Diagnostic Hook: record time to first token, tokens per second and total duration per call. Rising time to first token with steady tokens per second means queueing at the provider; falling tokens per second with steady time to first token on many concurrent streams points at your own event loop — check its lag before blaming the API.

Pitfalls & edge cases

  • httpx default timeouts. Measured: ReadTimeout at 5 s for a 6 s first token.
  • Catching only SDK exceptions. A mid-stream timeout escaped as httpx2.ReadTimeout.
  • Status 200 errors. Errors after the stream starts carry the original response's status.
  • asyncio.timeout inside an async generator. Measured: a slow consumer got CancelledError instead of TimeoutError.
  • Many streams on one loop. At 52–97 µs per token, parsing alone can saturate a core.

Frequently Asked Questions

How do I stream Claude responses with asyncio in Python?

Use AsyncAnthropic and async with client.messages.stream(...) as stream: async for text in stream.text_stream. Leaving the context closes the HTTP stream; call get_final_message() for usage and the stop reason.

Why does my LLM streaming request time out after 5 seconds?

httpx's default timeout is 5 s for reads, and the first token can take longer. The Anthropic SDK defaults to a 600 s read timeout; with raw httpx, raise the read timeout and bound the whole call with asyncio.timeout.

What exceptions can a streaming LLM call raise mid-stream?

In testing with anthropic 1.11.0, an error event after streaming began raised APIStatusError with status_code 200, and a per-request timeout during the stream raised httpx2.ReadTimeout, which is not an SDK exception. Catch both.

How much CPU does streaming LLM tokens use?

About 52–54 µs per token parsing SSE with httpx and 77–97 µs with the SDK's stream helper in testing — roughly half a core to a full core for 10,000 tokens per second across all streams.