Instrumenting httpx and aiohttp with OpenTelemetry¶
Outgoing HTTP calls are where most of an async service's latency goes, so they are the spans a trace most needs. The OpenTelemetry instrumentation packages for httpx and aiohttp create a client span for every request and inject the traceparent header so the downstream service joins the same trace. Tested with opentelemetry-sdk 1.45 and the 0.66b0 instrumentations against a local server: all 4,000 requests produced a client span and arrived with a traceparent header. The cost per request, measured with a batch span processor, was +427 µs for httpx (517 → 944 µs) and +94 µs for aiohttp (129 → 222 µs) — negligible for real network calls, noticeable for very chatty local traffic. One ordering trap: an aiohttp ClientSession created before instrument() was called produced no spans and no header; httpx clients created earlier were instrumented anyway. This guide sets up both clients, keeps cardinality and overhead in check, and verifies propagation.
Prerequisites¶
- Python 3.11+,
pip install opentelemetry-sdk opentelemetry-instrumentation-httpx opentelemetry-instrumentation-aiohttp-client. - Tracing setup, from tracing asyncio services with OpenTelemetry.
- Log correlation, from correlating logs and traces with request IDs.
1. Instrument at startup, before creating clients¶
Call instrument() once, as early as possible — before any client or session exists:
from opentelemetry.instrumentation.aiohttp_client import AioHttpClientInstrumentor
from opentelemetry.instrumentation.httpx import HTTPXClientInstrumentor
def setup_tracing() -> None:
configure_tracer_provider() # exporter, resource, sampler
HTTPXClientInstrumentor().instrument()
AioHttpClientInstrumentor().instrument()
@asynccontextmanager
async def lifespan(app):
setup_tracing() # first
app.state.http = httpx.AsyncClient() # then clients
app.state.session = aiohttp.ClientSession()
yield
...
Tested: an aiohttp session created before instrument() produced no spans and sent no traceparent, because aiohttp's instrumentation attaches trace configs when a session is created; a session created afterwards was fully instrumented. httpx clients created before and after both worked, because its instrumentation patches the transport at request time. Relying on that difference is fragile — instrument first, everywhere. With zero-code instrumentation (opentelemetry-instrument python app.py), this ordering is handled for you.
Verify: a test that creates clients the way production does finds a traceparent header on a request to a local server.
2. Check what the spans record¶
Client spans carry the method, URL and status code, and become children of whatever span is current when the request is made:
from opentelemetry import trace
tracer = trace.get_tracer(__name__)
async def get_profile(user_id: int) -> dict:
with tracer.start_as_current_span("get_profile") as span:
span.set_attribute("user.id", user_id)
response = await app.state.http.get(f"{PROFILE_URL}/users/{user_id}") # child span "GET"
response.raise_for_status()
return response.json()
Tested: spans were named by method ("GET") with http.method, http.url and http.status_code attributes; newer semantic-convention versions use names like http.request.method and url.full, selected with the OTEL_SEMCONV_STABILITY_OPT_IN environment variable. Because the span is a child of the current span, wrapping business operations in your own spans gives the trace readable structure — "get_profile" containing "GET" — instead of a flat list of HTTP calls.
Verify: a trace of one request shows your operation spans with HTTP client spans nested under them, and downstream services' server spans under those.
3. Keep URLs from exploding cardinality and leaking secrets¶
The full URL lands in a span attribute. That is fine for traces but dangerous if spans feed metrics, and URLs can carry tokens:
from opentelemetry.trace import Span
def request_hook(span: Span, request) -> None:
if span.is_recording():
span.set_attribute("peer.service", service_for_host(request.url.host))
def response_hook(span: Span, request, response) -> None:
if span.is_recording() and response.status_code >= 500:
span.set_attribute("error.upstream", True)
HTTPXClientInstrumentor().instrument(request_hook=request_hook, response_hook=response_hook)
# Never put credentials in query strings; if a legacy API requires it, redact before export:
# OTEL_INSTRUMENTATION_HTTP_CAPTURE_HEADERS_* settings control which headers are captured
Hooks add attributes such as the logical service name, which is what dashboards should group by — not the URL, which has an unbounded number of values (/users/1, /users/2, …). Query-string tokens and signed URLs (object storage presigned URLs, OAuth callbacks) end up verbatim in span data; prefer headers for credentials, and scrub URLs in a span processor where you cannot. Header capture is off by default and should stay that way for Authorization and cookies.
Verify: a search of exported spans for token=, signature= or X-Amz-Signature returns nothing.
4. Control the overhead¶
Every span costs CPU to create, record and export. For most services the measured +94 to +427 µs per request is a small fraction of a real network call; for services making thousands of local calls per request, it adds up:
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.sdk.trace.sampling import ParentBased, TraceIdRatioBased
provider = TracerProvider(sampler=ParentBased(TraceIdRatioBased(0.1))) # 10% of new traces
provider.add_span_processor(BatchSpanProcessor(exporter, max_queue_size=4096,
schedule_delay_millis=1000))
Use a batch processor (never the simple, synchronous one in production), and sample: unsampled spans are created as non-recording and cost much less. ParentBased keeps the decision consistent with the incoming request's, so a trace is either complete or absent. Sampling strategies for high-throughput services are in sampling traces in high-throughput async services. To exclude noisy endpoints — health checks of dependencies, metrics scrapes — use the instrumentors' URL exclusion settings.
Verify: CPU per request with tracing enabled at your sampling rate is within a few percent of tracing disabled, measured under load.
5. Verify propagation end to end¶
Spans in one service are only half the value; the trace must continue in the next. Test the header on the wire:
async def test_outgoing_requests_carry_traceparent(aiohttp_server_url):
received: list[str | None] = []
async def handler(request):
received.append(request.headers.get("traceparent"))
return web.Response(text="ok")
async with run_server(handler) as url:
with tracer.start_as_current_span("test"):
await app.state.http.get(url)
assert received[0] and received[0].startswith("00-") # W3C trace context
Tested in exactly this way, 4,000 of 4,000 requests carried the header. In production, check that downstream services' spans have the caller's span as parent; a trace that stops at the client span means the downstream service is not extracting context — or a proxy in between strips the header.
Verify: a trace for one request in the backend spans every service the request touched.
Verification¶
HTTP clients are traced well when:
- Instrumentation runs before any client or session is created.
- Client spans nest under operation spans, with service-level attributes for grouping.
- URLs carry no secrets, and header capture excludes credentials.
- Propagation is tested and visible across services in the backend.
Diagnostic Hook: count client spans per second against outgoing requests per second (from your own client metrics). A ratio well below the sampling rate means some clients are not instrumented — typically sessions created at import time, before setup ran.
Pitfalls & edge cases¶
- aiohttp sessions created before
instrument(). Tested: no spans, no header. - Synchronous span processors. Each request waits for export.
- URLs as metric dimensions. Unbounded cardinality.
- Secrets in query strings. They end up in span attributes.
Frequently Asked Questions¶
How do I add OpenTelemetry tracing to httpx?
Install opentelemetry-instrumentation-httpx and call HTTPXClientInstrumentor().instrument() at startup; every request then gets a client span and a traceparent header. In testing, 4,000 of 4,000 requests did.
Why are my aiohttp requests not traced?
The ClientSession was probably created before AioHttpClientInstrumentor().instrument() was called; in testing such a session produced no spans or headers. Instrument first, then create sessions.
How much overhead does OpenTelemetry add to HTTP requests?
In testing with a batch processor against a local server, about 427 µs per request for httpx and 94 µs for aiohttp — small compared with real network calls. Sampling reduces it further.
How do I keep sensitive URLs out of traces?
Do not put credentials in query strings; where unavoidable, scrub url attributes in a span processor before export, and leave header capture disabled for Authorization and cookies.
Related¶
- Observability & Tracing — up to the topic overview.
- Continuous profiling of async services — when spans show time but not where CPU went.
- Resilience, Cancellation & Error Handling — the section overview.