Streaming LLM Tokens Asynchronously in Python¶
Streaming is the default way to call a large language model from an interactive application: the first words arrive in a fraction of a second instead of after the whole answer is generated. From asyncio it is also the place where the usual HTTP assumptions break. A response can stay open for minutes, the first byte can take longer than a typical read timeout, errors can arrive after a 200 OK, and every token is a small event the event loop must parse. To measure these behaviours, the official anthropic Python SDK (1.11.0) and plain httpx were run on Python 3.14 against a local mock of the Messages API that streams server-sent events in the same format. With 50 concurrent streams of 400 tokens each, the client spent 52–54 µs of CPU per token parsing SSE with httpx and 77–97 µs through the SDK — about half a core to one core per 10,000 tokens per second. When the first token took 6 s, httpx's default timeout failed with ReadTimeout after 5.0 s while the SDK's 600 s read timeout let it finish in 6.5 s. A per-request timeout=3.0 that expired mid-stream raised httpx2.ReadTimeout, not the SDK's APITimeoutError. And an error event sent after streaming had begun surfaced as APIStatusError with status code 200. This guide streams tokens with those edges handled.
Prerequisites¶
- Python 3.11+,
pip install anthropic(tested with 1.11.0; it depends onhttpx2). - Streaming HTTP basics, from streaming large responses with httpx.
- The topic overview, Concurrent LLM API Calls.
1. Stream with the SDK's context manager¶
The SDK's messages.stream() helper opens the request, parses the server-sent events, and accumulates the final message for you:
import asyncio
from anthropic import AsyncAnthropic
client = AsyncAnthropic() # one client per process, reused
async def answer(prompt: str) -> str:
parts: list[str] = []
async with client.messages.stream(
model="claude-sonnet-5-5",
max_tokens=1024,
messages=[{"role": "user", "content": prompt}],
) as stream:
async for text in stream.text_stream:
parts.append(text) # or forward it to the user
final = await stream.get_final_message()
log.info("tokens in=%d out=%d stop=%s", final.usage.input_tokens,
final.usage.output_tokens, final.stop_reason)
return "".join(parts)
Against the mock, with 200 ms to the first token and 10 ms per token after it, the first text arrived after 222 ms and 200 tokens in 2.24 s; the final message reported 200 output tokens and the stop reason. The context manager is not decoration: leaving it — normally or by exception — closes the HTTP response, which is what tells the server to stop generating. The cancellation guide shows what happens without it.
Verify: the usage numbers from get_final_message() are logged for every call; usage is what token budgets and costs are computed from.
2. Budget CPU for the token rate¶
Every streamed chunk is an SSE event: bytes to decode, a line to split, JSON to parse, an object to build. On one event loop that adds up. Measured with 50 concurrent streams of 400 tokens at 5 ms per token — 20,000 tokens in about 2.3 s — two client implementations were compared:
# Raw httpx: parse SSE lines yourself
async with http.stream("POST", "/v1/messages", json=req | {"stream": True}, headers=headers) as r:
r.raise_for_status()
async for line in r.aiter_lines():
if line.startswith("data: "):
event = json.loads(line[6:])
if event["type"] == "content_block_delta":
handle(event["delta"]["text"])
The raw parser used 1.05–1.09 s of CPU (52–54 µs per token) with event-loop lag p99 of 2.4–3.3 ms. The SDK used 1.54–1.94 s (77–97 µs per token) with lag p99 of 4.2–12.2 ms across two runs — it builds typed event objects and an accumulated message snapshot. Neither is a problem for one user; at 10,000 tokens per second across all streams, the SDK alone would occupy most of a core. Services that relay many simultaneous streams should measure their own token rate, keep JSON work light per event, and run several worker processes, as in sizing uvicorn workers for async services.
Verify: at your peak concurrent streams, event-loop lag p99 stays below the latency your users notice.
3. Set timeouts for slow first tokens¶
A model may take seconds before the first token — long prompts, queued capacity, extended reasoning. A client timeout tuned for ordinary APIs kills those requests:
# httpx's default is 5 s for connect, read, write and pool
async with httpx.AsyncClient(base_url=URL) as http:
async with http.stream("POST", "/v1/messages", json=req) as r:
async for line in r.aiter_lines(): # ReadTimeout after 5.0 s if TTFT is 6 s
...
# The SDK's default: connect 5 s, read/write/pool 600 s
client = AsyncAnthropic() # completed the same request in 6.5 s
Measured against the mock with a 6-second time to first token: raw httpx with defaults raised ReadTimeout at 5.0 s; the SDK completed in 6.5 s. The SDK's read timeout is per read, not per request, so a stream that keeps producing tokens can run far longer than 600 s; to bound the whole call, put a deadline around it, which is the subject of the next step. Connect timeouts can stay short — 5 seconds is generous for establishing TLS to an API endpoint — and should, so an unreachable endpoint fails fast.
Verify: your read timeout exceeds the p99 time to first token observed in production, and a separate deadline bounds the total duration.
4. Catch the exceptions streaming actually raises¶
The SDK maps HTTP errors that happen before streaming starts to its own exception hierarchy. Failures after the response has started arrive differently, and two of them surprised in testing:
import httpx2 # the SDK's HTTP layer in anthropic 1.11.0
import anthropic
async def stream_text(prompt: str):
"""Yield text chunks; translate every failure into the application's own errors."""
try:
async with client.messages.stream(model=MODEL, max_tokens=1024,
messages=[{"role": "user", "content": prompt}]) as s:
async for text in s.text_stream:
yield text
except anthropic.APIStatusError as exc: # includes mid-stream error events
raise LLMError(exc.status_code, str(exc)) from exc
except (anthropic.APIConnectionError, httpx2.TransportError) as exc:
raise LLMUnavailable(str(exc)) from exc
async def answer(prompt: str, deadline: float = 120.0) -> str:
parts: list[str] = []
async with asyncio.timeout(deadline): # the deadline lives in the caller
async for text in stream_text(prompt):
parts.append(text)
return "".join(parts)
Measured: with the SDK's per-request timeout=3.0 and a 6-second gap mid-stream, the exception was httpx2.ReadTimeout, whose base classes are TimeoutException, TransportError, RequestError and HTTPError — it was neither anthropic.APITimeoutError nor anthropic.APIConnectionError, and an except anthropic.APIError would have let it escape. When the mock sent an error event with overloaded_error after 20 tokens, the SDK raised APIStatusError whose status_code was 200 — the status of the response that had already begun. asyncio.timeout(3.0) around the stream raised TimeoutError at exactly 3.0 s, independent of the SDK's internals. Where that timeout goes matters: placed inside an async generator around its yields, it misfired. With a consumer that awaited 50 ms per token, the deadline expired while the consumer — not the generator — was running, so the cancellation landed in the consumer's code and surfaced as a bare CancelledError after 5 tokens instead of a TimeoutError; with a fast consumer the same code raised TimeoutError as intended. Keep the deadline in the code that iterates, as above. The retry guide covers which of these to retry.
Verify: a test that injects a mid-stream timeout and a mid-stream error event gets your application's own exception types, not a raw transport error.
5. Reuse one client and keep streams independent¶
AsyncAnthropic owns an HTTP connection pool — up to 1,000 connections, 100 kept alive, by default — so create it once and share it across tasks. Each stream is an independent request on that pool, and concurrent streams do not block each other:
from contextlib import asynccontextmanager
@asynccontextmanager
async def lifespan(app):
app.state.llm = AsyncAnthropic(max_retries=2)
try:
yield
finally:
await app.state.llm.close() # closes pooled connections
Creating a client per request discards its pool and pays TLS setup on every call, the same cost measured for other HTTP clients in reusing aiohttp ClientSession across requests. The SDK's limits are generous; your concurrency limit should come from the provider's rate limits instead, as in fanning out LLM calls with bounded concurrency.
Verify: the number of client objects created over the life of the process is one per configuration, not one per request.
Verification¶
Async LLM streaming is production-ready when:
- Streams are consumed inside
messages.stream(), and usage is logged from the final message. - Read timeouts exceed the p99 time to first token, and
asyncio.timeoutbounds each call's total duration. - Mid-stream failures are caught, including
httpx2.TransportErrorandAPIStatusErrorwith status 200. - CPU per token is known, and event-loop lag stays acceptable at peak concurrent streams.
Diagnostic Hook: record time to first token, tokens per second and total duration per call. Rising time to first token with steady tokens per second means queueing at the provider; falling tokens per second with steady time to first token on many concurrent streams points at your own event loop — check its lag before blaming the API.
Pitfalls & edge cases¶
- httpx default timeouts. Measured:
ReadTimeoutat 5 s for a 6 s first token. - Catching only SDK exceptions. A mid-stream timeout escaped as
httpx2.ReadTimeout. - Status 200 errors. Errors after the stream starts carry the original response's status.
asyncio.timeoutinside an async generator. Measured: a slow consumer gotCancelledErrorinstead ofTimeoutError.- Many streams on one loop. At 52–97 µs per token, parsing alone can saturate a core.
Frequently Asked Questions¶
How do I stream Claude responses with asyncio in Python?
Use AsyncAnthropic and async with client.messages.stream(...) as stream: async for text in stream.text_stream. Leaving the context closes the HTTP stream; call get_final_message() for usage and the stop reason.
Why does my LLM streaming request time out after 5 seconds?
httpx's default timeout is 5 s for reads, and the first token can take longer. The Anthropic SDK defaults to a 600 s read timeout; with raw httpx, raise the read timeout and bound the whole call with asyncio.timeout.
What exceptions can a streaming LLM call raise mid-stream?
In testing with anthropic 1.11.0, an error event after streaming began raised APIStatusError with status_code 200, and a per-request timeout during the stream raised httpx2.ReadTimeout, which is not an SDK exception. Catch both.
How much CPU does streaming LLM tokens use?
About 52–54 µs per token parsing SSE with httpx and 77–97 µs with the SDK's stream helper in testing — roughly half a core to a full core for 10,000 tokens per second across all streams.
Related¶
- Concurrent LLM API Calls — up to the topic overview.
- Cancelling streaming LLM responses — making the server stop when you do.
- Concurrent Execution & Worker Patterns — the section overview.