Skip to content

Sampling Traces in High-Throughput Async Services

Tracing every request in a busy service costs CPU, network and storage, and most of those traces are identical successful requests nobody will look at. Sampling keeps a representative fraction plus the interesting ones. The CPU side is easy to underestimate. Measured with the OpenTelemetry SDK on an asyncio workload where each request created six spans (a request span and five database spans): with a no-op tracer, 19.9 µs of CPU per request; recording and exporting every trace, 194.1 µs — nearly ten times as much; sampling 10%, 56.1 µs; 1%, 40.5 µs; and with sampling turned off entirely but the SDK still installed, 33.3 µs, because unsampled spans are still created as lightweight non-recording objects. The exported span counts matched the rates — 120,000, 11,874, 1,284 and 0 from 20,000 requests. This guide chooses sampling that keeps the traces you need at a cost you can afford.

Prerequisites

1. Know what a sampled span costs

A recording span stores attributes, events and timing, then is queued, batched and serialized for export. A non-recording span is a small object that only carries context:

from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.sdk.trace.sampling import ParentBased, TraceIdRatioBased

provider = TracerProvider(sampler=ParentBased(TraceIdRatioBased(1.0)))   # record everything
provider.add_span_processor(BatchSpanProcessor(exporter))

Measured CPU per six-span request: 194.1 µs at 100% against 19.9 µs with no tracing — at 2,000 requests per second, that difference is about 35% of a CPU core spent on tracing. The cost scales with spans per request, so services that create a span per database row or per cache lookup pay it many times over. Attributes added to non-recording spans are discarded cheaply; guard expensive attribute computation with span.is_recording().

Verify: compare CPU per request with tracing at your planned rate against tracing disabled, under realistic load.

CPU per request (6 spans) by sampling rate 5 horizontal bars comparing no-op tracer (no SDK) with the others. CPU per request (6 spans) by sampling rate no-op tracer (no SDK) 19.9 µs sampled 100% 194.1 µs sampled 10% 56.1 µs sampled 1% 40.5 µs sampled 0% (ALWAYS_OFF) 33.3 µs opentelemetry-sdk 1.45, BatchSpanProcessor, 20,000 asyncio requests; exported 120,000 / 11,874 / 1,284 / 0 spans. Most of tracing's cost is recording and exporting; sampling removes it.

2. Use parent-based ratio sampling at the edge

Make the sampling decision once, when a trace starts, and have every service follow it:

from opentelemetry.sdk.trace.sampling import ParentBased, TraceIdRatioBased

sampler = ParentBased(
    root=TraceIdRatioBased(0.05),       # new traces: keep 5%
)                                        # traces started upstream: follow the caller's decision
provider = TracerProvider(sampler=sampler)

TraceIdRatioBased decides from the trace id itself, so every service that sees the same trace id with the same ratio makes the same decision, and ParentBased follows the incoming traceparent flag when there is one. The result is complete traces — never a trace with gaps where one service sampled and another did not. Measured, 10% sampling exported 11,874 of 120,000 spans (9.9%) and cut CPU per request from 194.1 to 56.1 µs. Set the rate from the request rate and how many traces per minute you actually need: at 2,000 requests per second, 1% is still 1,200 traces a minute.

Verify: in the tracing backend, sampled traces include spans from every service on the path.

3. Keep errors and slow requests with tail-based sampling

Head sampling decides before anyone knows whether the request will fail or be slow, so at 1% it keeps 1% of errors too. Tail-based sampling decides after the trace is complete, in a collector that buffers spans:

# OpenTelemetry Collector: keep all errors, all slow traces, and 2% of the rest
processors:
  tail_sampling:
    decision_wait: 10s
    policies:
      - name: errors
        type: status_code
        status_code: {status_codes: [ERROR]}
      - name: slow
        type: latency
        latency: {threshold_ms: 500}
      - name: baseline
        type: probabilistic
        probabilistic: {sampling_percentage: 2}

With tail sampling, services export every span (so their CPU cost is the 100% figure), and the collector keeps errors, slow traces and a baseline. That moves cost from storage to the service and the collector; a common compromise is head sampling at a moderate rate in services plus tail sampling in the collector to keep the interesting subset of those. Mark errors on spans (span.set_status(Status(StatusCode.ERROR)) or record_exception) so status-based policies can see them.

Verify: a deliberately failing request is found in the backend even when its trace id would not have been head-sampled.

Head versus tail sampling A grid of 3 rows by 4 columns. Head versus tail sampling approach decided service CPU keeps errors head, ParentBased ratio at trace start low (56 µs at 10%) at the sampled rate tail, in the collector after completion full (194 µs) all of them head + tail both moderate all errors among head-sampled Head sampling saves the service; tail sampling saves the interesting traces.

4. Sample within long or chatty traces

Some requests create hundreds of spans — a batch job, a request that queries per item. Even when the trace is sampled, recording every inner span may not be worth it:

tracer = trace.get_tracer(__name__)


async def process_batch(items) -> None:
    with tracer.start_as_current_span("process_batch") as batch_span:
        batch_span.set_attribute("batch.size", len(items))
        for i, item in enumerate(items):
            if i < 20 or i % 100 == 0:                         # first few and a sample of the rest
                with tracer.start_as_current_span("process_item") as span:
                    span.set_attribute("item.index", i)
                    await process(item)
            else:
                await process(item)                            # timed in aggregate, not per item
        batch_span.set_attribute("batch.items_processed", len(items))

Recording a sample of inner spans plus aggregate attributes on the parent keeps the trace readable and its cost bounded. Metrics are better than spans for per-item timing in large batches: a histogram of item durations costs far less than a span each. The same applies to hot inner loops — cache lookups, serialization — where a span per call can exceed the cost of the work.

Verify: the largest traces in the backend have a bounded span count.

5. Watch exporter queues and dropped spans

At high throughput, the batch processor's queue can fill faster than the exporter drains it, and spans are dropped silently:

processor = BatchSpanProcessor(
    exporter,
    max_queue_size=8192,            # spans buffered before dropping
    max_export_batch_size=1024,
    schedule_delay_millis=1000,
    export_timeout_millis=10000,
)

Dropped spans produce traces with missing pieces — worse than unsampled traces, because they look complete. The SDK logs a warning when the queue is full; monitor for it, and size the queue for your span rate times the export interval with headroom. If exports cannot keep up even with generous queues, the answer is a lower sampling rate or fewer spans per request, not a bigger buffer.

Verify: under peak load, no "queue is full" warnings appear and exported span counts match the sampling rate.

Which sampling setup fits this service? A decision on What does the traffic look like with 4 outcomes. Which sampling setup fits this service? What does the traffic look like? modest rate record 100% 194 µs/request is affordable high rate ParentBased ratio at edge e.g. 1-10% need every error/slow trace + tail sampling in collector export all spans huge traces sample inner spans, use metrics bounded size Decide the rate from traces needed per minute, not from habit.

Verification

Sampling is set up well when:

  • The decision is made once per trace with ParentBased ratio sampling.
  • Tracing CPU is measured at the chosen rate and is affordable.
  • Errors and slow requests are kept by tail sampling where needed.
  • Span counts per trace are bounded, and exporter queues never overflow.

Diagnostic Hook: export the number of spans started, sampled, exported and dropped per second (the SDK's own metrics or a custom span processor that counts). Exported much lower than sampled means the exporter cannot keep up; sampled much higher than your target rate means a service is not using ParentBased and is starting its own traces.

Pitfalls & edge cases

  • 100% sampling at high throughput. Measured: nearly 10× the CPU of untraced requests.
  • Independent per-service sampling. Traces with gaps; use ParentBased.
  • Head sampling alone for errors. At 1%, 99% of failures have no trace.
  • Unbounded spans per trace. Batch jobs produce traces nobody can read.

Frequently Asked Questions

How much CPU does OpenTelemetry tracing use in Python?

In testing, a six-span request cost 194 µs of CPU when every trace was recorded and exported, against 20 µs untraced; 10% sampling brought it to 56 µs.

What sampler should I use for OpenTelemetry in Python?

ParentBased(TraceIdRatioBased(rate)) for most services: new traces are sampled by trace id at the chosen rate, and downstream services follow the caller's decision.

How do I keep all error traces when sampling?

Use tail-based sampling in the OpenTelemetry Collector with a status-code policy; services export all spans and the collector keeps errors, slow traces and a baseline fraction.

Does disabling sampling remove tracing overhead completely?

Not entirely: with ALWAYS_OFF but the SDK installed, requests still cost 33 µs against 20 µs with a no-op tracer, because non-recording spans are still created.