Sampling Traces in High-Throughput Async Services¶
Tracing every request in a busy service costs CPU, network and storage, and most of those traces are identical successful requests nobody will look at. Sampling keeps a representative fraction plus the interesting ones. The CPU side is easy to underestimate. Measured with the OpenTelemetry SDK on an asyncio workload where each request created six spans (a request span and five database spans): with a no-op tracer, 19.9 µs of CPU per request; recording and exporting every trace, 194.1 µs — nearly ten times as much; sampling 10%, 56.1 µs; 1%, 40.5 µs; and with sampling turned off entirely but the SDK still installed, 33.3 µs, because unsampled spans are still created as lightweight non-recording objects. The exported span counts matched the rates — 120,000, 11,874, 1,284 and 0 from 20,000 requests. This guide chooses sampling that keeps the traces you need at a cost you can afford.
Prerequisites¶
- Python 3.11+,
pip install opentelemetry-sdk(measured with 1.45). - Tracing setup, from tracing asyncio services with OpenTelemetry.
- A collector if you want tail-based sampling (the OpenTelemetry Collector).
1. Know what a sampled span costs¶
A recording span stores attributes, events and timing, then is queued, batched and serialized for export. A non-recording span is a small object that only carries context:
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.sdk.trace.sampling import ParentBased, TraceIdRatioBased
provider = TracerProvider(sampler=ParentBased(TraceIdRatioBased(1.0))) # record everything
provider.add_span_processor(BatchSpanProcessor(exporter))
Measured CPU per six-span request: 194.1 µs at 100% against 19.9 µs with no tracing — at 2,000 requests per second, that difference is about 35% of a CPU core spent on tracing. The cost scales with spans per request, so services that create a span per database row or per cache lookup pay it many times over. Attributes added to non-recording spans are discarded cheaply; guard expensive attribute computation with span.is_recording().
Verify: compare CPU per request with tracing at your planned rate against tracing disabled, under realistic load.
2. Use parent-based ratio sampling at the edge¶
Make the sampling decision once, when a trace starts, and have every service follow it:
from opentelemetry.sdk.trace.sampling import ParentBased, TraceIdRatioBased
sampler = ParentBased(
root=TraceIdRatioBased(0.05), # new traces: keep 5%
) # traces started upstream: follow the caller's decision
provider = TracerProvider(sampler=sampler)
TraceIdRatioBased decides from the trace id itself, so every service that sees the same trace id with the same ratio makes the same decision, and ParentBased follows the incoming traceparent flag when there is one. The result is complete traces — never a trace with gaps where one service sampled and another did not. Measured, 10% sampling exported 11,874 of 120,000 spans (9.9%) and cut CPU per request from 194.1 to 56.1 µs. Set the rate from the request rate and how many traces per minute you actually need: at 2,000 requests per second, 1% is still 1,200 traces a minute.
Verify: in the tracing backend, sampled traces include spans from every service on the path.
3. Keep errors and slow requests with tail-based sampling¶
Head sampling decides before anyone knows whether the request will fail or be slow, so at 1% it keeps 1% of errors too. Tail-based sampling decides after the trace is complete, in a collector that buffers spans:
# OpenTelemetry Collector: keep all errors, all slow traces, and 2% of the rest
processors:
tail_sampling:
decision_wait: 10s
policies:
- name: errors
type: status_code
status_code: {status_codes: [ERROR]}
- name: slow
type: latency
latency: {threshold_ms: 500}
- name: baseline
type: probabilistic
probabilistic: {sampling_percentage: 2}
With tail sampling, services export every span (so their CPU cost is the 100% figure), and the collector keeps errors, slow traces and a baseline. That moves cost from storage to the service and the collector; a common compromise is head sampling at a moderate rate in services plus tail sampling in the collector to keep the interesting subset of those. Mark errors on spans (span.set_status(Status(StatusCode.ERROR)) or record_exception) so status-based policies can see them.
Verify: a deliberately failing request is found in the backend even when its trace id would not have been head-sampled.
4. Sample within long or chatty traces¶
Some requests create hundreds of spans — a batch job, a request that queries per item. Even when the trace is sampled, recording every inner span may not be worth it:
tracer = trace.get_tracer(__name__)
async def process_batch(items) -> None:
with tracer.start_as_current_span("process_batch") as batch_span:
batch_span.set_attribute("batch.size", len(items))
for i, item in enumerate(items):
if i < 20 or i % 100 == 0: # first few and a sample of the rest
with tracer.start_as_current_span("process_item") as span:
span.set_attribute("item.index", i)
await process(item)
else:
await process(item) # timed in aggregate, not per item
batch_span.set_attribute("batch.items_processed", len(items))
Recording a sample of inner spans plus aggregate attributes on the parent keeps the trace readable and its cost bounded. Metrics are better than spans for per-item timing in large batches: a histogram of item durations costs far less than a span each. The same applies to hot inner loops — cache lookups, serialization — where a span per call can exceed the cost of the work.
Verify: the largest traces in the backend have a bounded span count.
5. Watch exporter queues and dropped spans¶
At high throughput, the batch processor's queue can fill faster than the exporter drains it, and spans are dropped silently:
processor = BatchSpanProcessor(
exporter,
max_queue_size=8192, # spans buffered before dropping
max_export_batch_size=1024,
schedule_delay_millis=1000,
export_timeout_millis=10000,
)
Dropped spans produce traces with missing pieces — worse than unsampled traces, because they look complete. The SDK logs a warning when the queue is full; monitor for it, and size the queue for your span rate times the export interval with headroom. If exports cannot keep up even with generous queues, the answer is a lower sampling rate or fewer spans per request, not a bigger buffer.
Verify: under peak load, no "queue is full" warnings appear and exported span counts match the sampling rate.
Verification¶
Sampling is set up well when:
- The decision is made once per trace with
ParentBasedratio sampling. - Tracing CPU is measured at the chosen rate and is affordable.
- Errors and slow requests are kept by tail sampling where needed.
- Span counts per trace are bounded, and exporter queues never overflow.
Diagnostic Hook: export the number of spans started, sampled, exported and dropped per second (the SDK's own metrics or a custom span processor that counts). Exported much lower than sampled means the exporter cannot keep up; sampled much higher than your target rate means a service is not using ParentBased and is starting its own traces.
Pitfalls & edge cases¶
- 100% sampling at high throughput. Measured: nearly 10× the CPU of untraced requests.
- Independent per-service sampling. Traces with gaps; use
ParentBased. - Head sampling alone for errors. At 1%, 99% of failures have no trace.
- Unbounded spans per trace. Batch jobs produce traces nobody can read.
Frequently Asked Questions¶
How much CPU does OpenTelemetry tracing use in Python?
In testing, a six-span request cost 194 µs of CPU when every trace was recorded and exported, against 20 µs untraced; 10% sampling brought it to 56 µs.
What sampler should I use for OpenTelemetry in Python?
ParentBased(TraceIdRatioBased(rate)) for most services: new traces are sampled by trace id at the chosen rate, and downstream services follow the caller's decision.
How do I keep all error traces when sampling?
Use tail-based sampling in the OpenTelemetry Collector with a status-code policy; services export all spans and the collector keeps errors, slow traces and a baseline fraction.
Does disabling sampling remove tracing overhead completely?
Not entirely: with ALWAYS_OFF but the SDK installed, requests still cost 33 µs against 20 µs with a no-op tracer, because non-recording spans are still created.
Related¶
- Observability & Tracing — up to the topic overview.
- Instrumenting httpx and aiohttp with OpenTelemetry — where many of a service's spans come from.
- Resilience, Cancellation & Error Handling — the section overview.