OpenTelemetry Tail Sampling: Cutting Trace Costs Without Losing the Slow and Broken Requests
Most teams discover trace sampling the same way: the observability bill arrives, someone sets a 5 percent sampler in the SDK, and the bill drops. Then an incident happens, and the one trace that would explain it was in the 95 percent that were thrown away before anyone knew it was interesting. That is the structural flaw of deciding at the start of a request what to keep. OpenTelemetry tail sampling reverses the order: it buffers a whole trace, waits until the outcome is known, and only then decides whether the trace deserves storage.
This matters now because trace volume grows with every microservice, every retry, and every LLM or agent call you instrument, while the share of traces anyone ever opens stays tiny. Keeping all of them is expensive. Keeping a random few is cheap but blind. Keeping the right few requires a stateful tier in your collector fleet, and that tier has real failure modes.
You will leave with a working two-tier collector design, runnable YAML, a memory-sizing method for the buffer, a policy order that protects errors and latency outliers, and a worked cost example with clearly labelled illustrative numbers.
What this covers: head versus tail sampling, the tail_sampling processor and its policies, routing by trace ID with the load-balancing exporter, consistent probability sampling and W3C tracestate, computing RED metrics before sampling, sizing, failure modes, and a rollout checklist.
Context and Background
Distributed tracing records a request as a tree of spans that share a trace ID. Every span carries timing, status, and attributes, which makes a trace the highest-fidelity record of what happened to a single request. It is also the most voluminous telemetry signal. A single checkout call that fans out to twenty services, each with a database client span and an HTTP client span, easily produces dozens of spans. Multiply that by thousands of requests per second and the arithmetic becomes the largest line item in many observability budgets.
The OpenTelemetry project documents sampling as one of the most effective ways to reduce observability cost without losing visibility, and it names filtering and aggregation as the alternatives. Its sampling concepts page splits the field into two families. Head sampling decides as early as possible, usually from the trace ID and a target percentage, without inspecting the trace as a whole. Tail sampling decides after considering all or most of the spans in the trace, so it can act on criteria that only exist later: an error status three hops downstream, or a total duration that no single span reveals.
The same documentation is candid about the cost. Tail sampling is not “set and forget”, and the components implementing it must be stateful systems that accept and store a large amount of data. It adds that tail samplers have often ended up as vendor-specific technology. The Collector’s tail_sampling processor, which lives in the contrib repository and is also included in the Kubernetes distribution, is the open implementation most self-hosted teams use. Its traces pipeline support is currently documented as beta, which is worth remembering when you plan upgrades.
If you are new to how the Collector is assembled from receivers, processors, exporters, and connectors, start with our guide to OpenTelemetry Collector architecture and pipelines. If your fleet runs Grafana’s distribution, the Grafana Alloy OpenTelemetry Collector tutorial covers how the same components are expressed there.
The thesis of this post is narrower than “use tail sampling”. It is that tail sampling is a data-placement problem before it is a policy problem. Writing the policy YAML takes ten minutes. Getting every span of every trace to the same process, with enough memory and a long enough wait, and without losing your service-level metrics in the process, is where the engineering effort goes. Teams that treat it as a config change get surprised. Teams that treat it as a small stateful service get a stable pipeline.
Head Sampling Versus Tail Sampling: The Core Trade
Head sampling is a decision made once, at the start of the trace, from information available at that moment. Tail sampling is a decision made once, at the end, from the full evidence. In one sentence: head sampling is cheap and blind, tail sampling is informed and stateful, and most production systems that must sample at scale end up using both in sequence.

Figure 1: Where each strategy decides. Head sampling commits at the first span, before the outcome exists. Tail sampling buffers spans and decides once the trace is complete or the wait expires.
Figure 1 shows the asymmetry. On the head-sampling path, the decision happens at the moment the root span starts. At that moment the system cannot know whether the request will fail, stall on a lock, or hit a cold cache. The sampler can only apply a probability, so it keeps a uniform slice of traffic. That slice is statistically representative, which is exactly what you want for capacity planning and exactly what you do not want for debugging, because the rare event you care about is rare in the sample too.
What head sampling does well
Head sampling is efficient and easy to reason about. When a sampler such as the SDK’s trace-ID-ratio based sampler drops a trace, the application never records the child spans at all, so you save CPU, memory, and network in the instrumented service, not only storage downstream. It also needs no coordination: each service reaches the same verdict because the decision is derived from the trace ID and propagated in the trace context flags. There is no buffer to size and no tier to scale.
The limitation is spelled out in the OpenTelemetry documentation: it is not possible to make a sampling decision based on data in the entire trace. You cannot promise to keep every trace containing an error, because at the head there is no error yet. If your error rate is 0.5 percent and you head-sample at 5 percent, you keep roughly 0.025 percent of all requests as error traces, which is one in four thousand. Many incidents involve a failure mode that is rarer than that.
What tail sampling adds
Tail sampling lets the policy see the answer. The documented use cases are the ones operators actually ask for: always keep traces that contain an error, keep traces whose overall latency exceeds a threshold, keep traces based on attribute values such as sampling more traffic from a newly deployed service, and apply different rates to low-volume and high-volume services. Each of these is a function of the finished trace, so none can be implemented at the head.
The price is that the sampler now holds data. Every span that arrives is parked in memory, indexed by trace ID, until the decision wait elapses. The processor is stateful by construction, which leads to the central constraint of this whole topic: all spans of a trace must reach the same collector instance. A tail sampler that sees only half of a trace will evaluate the wrong evidence. It might see the fast half and discard a slow trace, or see no error span and drop a failed request.
Why most production systems use both
The OpenTelemetry documentation notes that tail sampling is sometimes combined with head sampling, where head sampling first trims a very large volume to protect the pipeline. That is a sensible pattern when your raw span rate would otherwise overwhelm the tail tier. The cost is that anything the head sampler discards is invisible to the tail sampler, so the head rate must be high enough that the rare events you care about still survive in expectation. For most teams, the first step is to avoid head sampling in the SDK entirely (use parent-based always-on), do the heavy lifting in a collector tier, and add head sampling only when measurements say the tail tier cannot keep up.
There is a second subtlety. Head sampling in the SDK reduces the load on the application, whereas tail sampling does not: the application still creates and exports every span. Your services pay the instrumentation overhead for 100 percent of traces and the saving appears only after the collector. For latency-sensitive services that overhead is the reason some teams keep a modest head rate. It is a judgement call, and the right way to make it is to measure exporter overhead in a representative service rather than to assume it is negligible.
The Reference Architecture: Two Tiers, Routed by Trace ID
The processor holds spans in memory until a decision is made, so a tail-sampling deployment is a two-layer system. The first layer receives spans from applications and forwards each span to a specific second-layer replica chosen from the trace ID. The second layer runs the sampling policies.

Figure 2: Tier 1 gateways hash the trace ID onto a ring of Tier 2 replicas, so every span of a given trace lands on the same tail sampler. The resolver keeps the ring in sync with the replica set.
The Collector contrib documentation for the tail sampling processor recommends exactly this structure: a front layer running the load-balancing exporter and a back layer running the tail sampling processor. It notes that running both pipelines on one instance is technically possible but that separate layers give better failure isolation. The reason is operational. Tier 1 is stateless and cheap, so you can restart it freely. Tier 2 holds in-flight traces, so a restart discards everything buffered. Keeping them apart means a Tier 1 rollout never costs you traces.
The load-balancing exporter
The load_balancing exporter (older configurations spell it loadbalancing, which still works as a deprecated alias) creates one OTLP sub-exporter per backend and routes by a chosen key. For traces the default routing_key is traceID, and for logs and metrics the default is service. The documented keys include service, traceID, resource, metric, streamID, and attributes, with traceID valid for spans and logs but not metrics. For tail sampling you want traceID, which maps every span of one trace to one backend.
Backends are discovered through a resolver, and exactly one may be configured. The static resolver takes a fixed list of hostnames. The DNS resolver re-queries a hostname periodically, with documented defaults of port 4317, a 5-second interval, and a 1-second timeout. The Kubernetes resolver watches the Service’s EndpointSlices and reacts faster than DNS. It needs RBAC permission to get, list, and watch discovery.k8s.io/v1 EndpointSlices, and without that permission the resolver cache stays empty, which looks like “no backends” and drops data. An AWS Cloud Map resolver also exists, returning at most 100 hosts.
A headless Kubernetes Service with the Kubernetes resolver is the usual choice, because DNS caching can leave Tier 1 routing to a pod that no longer exists. The documentation says plainly that topology changes take time to reach the exporter and that the Kubernetes resolver propagates them faster than DNS.
Consistent hashing and what a scale event does
The exporter assigns a trace to a backend with a consistent hash over the current endpoint list. When Tier 2 scales from three replicas to four, a fraction of trace IDs remap to a different replica. Traces that were mid-flight when the ring changed will have their early spans in the old replica and their later spans in the new one. Each replica then holds a partial trace and evaluates it on partial evidence.
This is the single most underappreciated failure mode, and it is inherent rather than a bug. You cannot make it zero; you can make it rare and survivable by scaling Tier 2 deliberately, not on every CPU blip, and by using conservative policies that default to keep when the evidence looks incomplete. We return to this in the failure-modes section.
Resiliency settings you should turn on
Each sub-exporter has its own queue, retry, and timeout settings, and by default the exporter does not reroute data to a healthy backend when a delivery fails. There are two layers of settings. The otlp section governs short-term backend trouble. The load-balancer-level timeout, retry_on_failure, and sending_queue settings, which are off by default, help when the endpoint list changes often, as in Kubernetes. If Tier 2 pods churn, enable them; otherwise a deployment can lose spans during a rollout. The documentation also states that persistent queues are not supported at the sub-exporter level because all sub-exporters share one queue configuration.
Running several Tier 1 collectors with identical configuration gives consistent routing and removes a single point of failure, since every Tier 1 instance computes the same trace-to-backend mapping from the same endpoint list.
Walk-Through: Policies, Sizing, and a Runnable Pipeline
This section builds the pipeline in the order you would actually deploy it: the Tier 1 gateway, the Tier 2 sampler with policies, the sizing of the buffer, and the service-metrics branch that must run before any sampling.
Tier 1: the routing gateway
The gateway does no sampling. It receives OTLP, protects itself with the memory limiter, and forwards by trace ID. The configuration below is a minimal sketch using the Kubernetes resolver. Names such as otel-sampler-headless are placeholders for your own Service.
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
processors:
memory_limiter:
check_interval: 1s
limit_percentage: 80
spike_limit_percentage: 20
exporters:
load_balancing:
routing_key: traceID
protocol:
otlp:
tls:
insecure: true
timeout: 5s
retry_on_failure:
enabled: true
sending_queue:
enabled: true
queue_size: 5000
resolver:
k8s:
service: otel-sampler-headless.observability
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter]
exporters: [load_balancing]
Field names such as the resolver block and the queue keys follow the contrib README at the time of writing; always diff against the README for your exact Collector version, because this component’s configuration schema has been renamed and extended more than once.
Tier 2: the tail sampler
The sampler receives from Tier 1, evaluates policies, and exports survivors. The configuration below combines the policy types most teams need. It is illustrative in its thresholds, which you must tune to your own latency distribution.
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
processors:
memory_limiter:
check_interval: 1s
limit_percentage: 80
spike_limit_percentage: 20
tail_sampling:
decision_wait: 30s
num_traces: 200000
expected_new_traces_per_sec: 1000
decision_cache:
sampled_cache_size: 200000
non_sampled_cache_size: 200000
policies:
- name: keep-errors
type: status_code
status_code:
status_codes: [ERROR]
- name: keep-slow
type: latency
latency:
threshold_ms: 1500
- name: checkout-extra
type: and
and:
and_sub_policy:
- name: is-checkout
type: string_attribute
string_attribute:
key: http.route
values: ["/api/checkout"]
- name: checkout-quarter
type: probabilistic
probabilistic:
sampling_percentage: 25
- name: baseline
type: probabilistic
probabilistic:
sampling_percentage: 2
- name: safety-valve
type: rate_limiting
rate_limiting:
spans_per_second: 50000
batch:
send_batch_size: 8192
timeout: 5s
exporters:
otlp/backend:
endpoint: traces-backend.observability:4317
tls:
insecure: true
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, tail_sampling, batch]
exporters: [otlp/backend]
Two points about this file. First, the policy list is an OR: a trace is kept if any listed policy votes to sample it. Second, the rate_limiting entry listed that way is a policy that votes “sample” while under budget, which is not the same as a cap on the whole pipeline. If you want a hard ceiling across everything, use a composite policy, covered below. Treat the sketch above as a starting shape, not a drop-in.
The policy vocabulary
The processor documents a broad set of policy types. The ones that matter for cost control are these.
status_code matches traces containing a span with the listed status codes (OK, ERROR, UNSET). It is the cheapest and highest-value policy: it ensures failures are never sampled away. A caveat that surprises people is that status is whatever instrumentation set. An HTTP 500 returned by a server span is marked as an error by convention, but an application that catches an exception and returns a friendly 200 with a failure payload will not trip this policy. Teams that care add a string_attribute or ottl_condition policy keyed on their own business-failure attribute.
latency keeps traces whose duration, measured from the earliest span start to the latest span end, exceeds threshold_ms, with an optional upper_threshold_ms. Note that it measures the trace, not the root span, so a trace with a long-running background child can look slow even when the user-facing request was fast. Choose thresholds from your measured latency distribution, per route if routes differ greatly. A single global threshold is either too low for batch endpoints or too high for interactive ones.
probabilistic samples a configured percentage of traces, with an optional hash_salt. It is your unbiased baseline: the slice of healthy traffic you keep to learn what normal looks like. Without a baseline you can tell what failed but not how failure compares with success.
string_attribute, numeric_attribute, and boolean_attribute match span or resource attributes, with options such as regex matching and invert_match. trace_state matches a key in the W3C tracestate. span_count keeps traces with a span count in a range, which is a cheap way to find pathological fan-out. ottl_condition evaluates OpenTelemetry Transformation Language conditions against spans and span events for cases the typed policies cannot express.
Two cost-protection policies are worth knowing. rate_limiting takes spans_per_second, and bytes_limiting takes bytes_per_second with a burst capacity that defaults to twice the rate. Both bound what a policy can admit. Logical combinators and, not, and drop compose policies, and composite allocates a shared budget across sub-policies using max_total_spans_per_second, a policy_order, and a rate_allocation that assigns percentages.

Figure 3: Policy evaluation as a decision flow. It shows a conceptual order: errors first, then latency, then business-critical routes, then a baseline probability, then a rate guard.
The diagram expresses intent. In the actual processor, policies are evaluated against the buffered trace and combined by OR semantics, so ordering does not change which traces are kept in the simple list form. Order matters for the composite policy, where policy_order and the allocation determine who gets budget first, and it matters for the sample_on_first_match option, which short-circuits evaluation to save CPU on expensive policy sets.
Sizing the buffer: decision_wait and num_traces
decision_wait is how long the processor waits, counted from the first span of a trace it sees, before it evaluates policies. The documented default is 30 seconds. num_traces is the maximum number of traces held in memory, with a documented default of 50,000. expected_new_traces_per_sec pre-allocates internal data structures, and its default of 0 means no pre-allocation.
The relationship is Little’s law. The number of traces in the buffer at steady state is approximately the arrival rate of new traces multiplied by the wait. At 2,000 new traces per second and a 30-second wait, that is 60,000 traces in flight. The default num_traces of 50,000 would already be too small. When the buffer is full, the processor evicts the oldest trace, which means it is decided early on incomplete evidence, and the processor reports this through the otelcol_processor_tail_sampling_sampling_trace_dropped_too_early metric. Alert on that metric being non-zero.
The more important question is how to choose the wait. It must exceed the longest trace you care about, because the processor decides when the timer fires, not when the trace is “done”. The trace-completeness gap is why teams pick a wait from the tail of their trace duration distribution, add margin, and accept that very long traces such as asynchronous workflows will be decided in fragments. Longer waits cost memory linearly; shorter waits cost decision quality. There is no free setting.
The processor also exposes decision_cache, with sampled_cache_size and non_sampled_cache_size both defaulting to 0 (inactive). The cache remembers recent decisions by trace ID so that late-arriving spans of an already-decided trace follow the same verdict instead of starting a new buffer entry. The otelcol_processor_tail_sampling_sampling_late_span_age histogram shows how late stragglers arrive and is the right input for tuning both the cache and the wait.
Recent versions add related knobs documented in the README: decision_wait_after_root_received (disabled by default) to decide sooner once a root span arrives, maximum_trace_size_bytes to drop oversized traces immediately, and num_shards (default 1, maximum 256) for parallel evaluation. Check which of these exist in the version you deploy.
Memory arithmetic with explicit assumptions
Memory use depends on spans, not traces. Suppose, purely as an illustration, 2,000 new traces per second, 40 spans per trace on average, and 2 KB of resident memory per buffered span after the Collector’s internal representation and attributes. These are assumptions to be replaced by your own measurements.
In flight: 2,000 × 30 = 60,000 traces. Spans: 60,000 × 40 = 2.4 million. Memory: 2.4 million × 2 KB ≈ 4.8 GB across the whole Tier 2. With four replicas, that is about 1.2 GB per replica for the buffer alone, before the batch queue, exporter queue, and Go runtime overhead. With a 50 percent headroom target you provision roughly 2 GB limits per replica and set num_traces to at least the per-replica share of 60,000 plus slack, so about 20,000 to 30,000 per replica in this example, not the global figure. Remember that num_traces is per processor instance.
The practical method is to measure: run one replica with production-like traffic for an hour, record resident memory and num_traces utilisation, and extrapolate. Spans from AI agent workloads or deep batch jobs can be an order of magnitude larger than the average web request, which is why per-span byte size, not trace count, is the number to watch.
Service metrics before sampling
Here is a trap that bites teams a month after rollout. Once traces are sampled, any request-rate or error-rate computed from the stored traces is wrong, because you stored 4 percent of traffic and error traces are over-represented by design. Dashboards that count traces will show a collapsed request rate and an inflated error ratio.
The fix is to compute RED metrics (rate, errors, duration) from 100 percent of spans before the sampler runs. The span_metrics connector (older configuration used the name spanmetrics, now deprecated) consumes spans in a traces pipeline and emits metrics into a metrics pipeline. It produces traces.span.metrics.calls and traces.span.metrics.duration, with a default explicit histogram, a default 60-second flush interval, and cumulative temporality unless you change it. Its stability is documented as alpha, and the documentation warns that the default duration unit is milliseconds today but is planned to change to seconds behind a feature gate, so set histogram.unit explicitly.

Figure 4: Pipeline layout. The span_metrics connector taps the full stream before the sampler, so RED metrics reflect 100 percent of traffic while only a small share is stored as traces.
In the Collector configuration, the connector appears as an exporter of the traces pipeline and a receiver of a metrics pipeline.
connectors:
span_metrics:
histogram:
unit: ms
dimensions:
- name: http.route
- name: http.request.method
metrics_flush_interval: 60s
aggregation_cardinality_limit: 5000
service:
pipelines:
traces/in:
receivers: [otlp]
processors: [memory_limiter]
exporters: [span_metrics, load_balancing]
metrics/spanmetrics:
receivers: [span_metrics]
exporters: [prometheusremotewrite]
This fragment shows the idea on the Tier 1 gateway: the same traces pipeline fans out to the connector and to the load balancer, so metrics are computed before any trace is dropped. Where you place the connector, Tier 1 or Tier 2, has consequences. On Tier 1 each gateway sees only a share of the traffic and emits partial series, which a metrics backend can sum. On Tier 2 each replica sees whole traces but only the traces routed to it. Either works if you aggregate downstream. What does not work is computing metrics after the sampler.
Mind cardinality. Every dimension multiplies series count. Putting a raw URL path or user ID in dimensions turns a cheap metric into an expensive one, which defeats the cost goal. The aggregation_cardinality_limit field exists for exactly this reason.
If you use exemplars, you can link a metric data point to a trace ID, but the linked trace may have been dropped by the sampler. Link exemplars only for traces you know you kept, or accept that some links will dangle.
Consistent Probability Sampling and W3C Tracestate
Tail sampling and probability sampling answer different questions, and mixing them without a shared encoding produces statistics you cannot trust. The tail sampler tells you which traces to keep. A consistent probability scheme tells every downstream consumer how many original traces each kept trace represents.
The OpenTelemetry specification for probability sampling in tracestate is, as of this writing, marked “Development” on the specification site, so treat the details below as the current design rather than a frozen contract. It defines two keys inside the ot vendor section of the W3C tracestate header. The key th carries a rejection threshold, and rv carries an explicit 56-bit randomness value.
The threshold T is derived from the sampling probability as T = (1 - p) × 2^56, encoded in hexadecimal with trailing zeros removed. The specification’s own example is 1 percent sampling, which encodes as ot=th:fd70a4. A threshold of zero means everything is kept. A span is kept when its randomness value R is greater than or equal to T. If rv is absent, the randomness comes from the least significant 56 bits of the trace ID, as in the W3C Trace Context Level 2 draft.
Why consistency matters for sampling in tiers
The specification defines consistent sampling this way: if a span is kept with probability p1, any span in the same trace sampled with probability p2 greater than or equal to p1 is also kept. That gives you coherent traces across heterogeneous samplers. If your edge gateway samples at 50 percent and your inner services sample at 100 percent, no inner span is dropped that the edge kept.
The adjusted count is the reciprocal of the sampling probability, which the specification expresses as 2^56 / (2^56 - T). If a trace was kept at 2 percent, it stands in for 50 original traces. A backend that knows the threshold can multiply counts by the adjusted count and recover an unbiased estimate of true volume, even though you stored a fraction of it. A backend that does not know the threshold can only count stored traces.
What the Collector provides today
The probabilistic sampler processor implements this scheme for traces and is documented as beta for traces and alpha for logs. It has three modes. The default hash_seed mode applies an FNV hash to the trace ID and uses 14 bits of information. The documentation requires all collectors in the same tier to share one hash_seed, because otherwise they disagree about which traces to keep and sub-trace completeness breaks. The proportional mode uses 56 bits of randomness and samples each item independently of earlier samplers. The equalizing mode accounts for prior sampling and lowers items to a uniform minimum probability. The documentation recommends proportional for traces because it follows the OpenTelemetry and W3C specifications, and notes that sampling_precision defaults to 4 hexadecimal digits for the threshold.
The tail sampling processor has an alpha feature gate named usetracestate that reads and writes these probability fields in tracestate. It is off by default, and the README says it should not be combined with sample_on_first_match. I could not verify from the documentation I read how the processor merges a tail-sampling decision into an upstream threshold, so I will not describe the exact arithmetic. If you need statistically correct counts from traces that passed through both a probability sampler and a tail sampler, test the resulting th values on a staging pipeline before building dashboards on them.
A practical stance
Not every team needs the tracestate machinery. If you store both traces and span-derived metrics, and compute RED metrics before sampling, you already have unbiased volume numbers from the metrics and a biased, deliberately informative sample from the traces. That is a coherent design. Reach for consistent probability sampling when you need to answer population questions, such as “what fraction of all requests used feature X”, directly from trace data, or when several sampling tiers must agree on completeness. For most cost-control projects the simpler design is enough.
Worked Cost Arithmetic (Illustrative)
The numbers below are invented to show the mechanics. Substitute your own traffic, span size, and backend price. Nothing here is a benchmark or a vendor quote.
Assumptions. A platform handles 2,000 new traces per second. The average trace has 40 spans, and the average span is 800 bytes on the wire to the backend. The backend charges a flat 0.30 USD per GB ingested. Of all traces, 1.5 percent contain an error and 1 percent exceed the latency threshold, with the overlap such that the union of “error or slow” is 2.2 percent.
Raw volume. Spans per second: 2,000 × 40 = 80,000. Bytes per second: 80,000 × 800 = 64 MB/s. Per day: 64 MB × 86,400 ≈ 5.5 TB, which is about 5,530 GB. At 0.30 USD per GB that is roughly 1,660 USD per day, or about 50,000 USD per month.
Head sampling at 5 percent. Volume falls to 5 percent, roughly 277 GB per day and 83 USD per day, about 2,500 USD per month. The error-trace retention is 5 percent of 1.5 percent, or 0.075 percent of all traces, so 95 percent of failing requests have no trace.
Tail sampling. Keep 100 percent of the error-or-slow union (2.2 percent of traces) and 2 percent of the remaining 97.8 percent, which is 1.956 percent. The total kept fraction is 4.156 percent, so about 230 GB per day. At 0.30 USD per GB that is about 69 USD per day, or roughly 2,070 USD per month. You kept 100 percent of failing and slow requests, not 5 percent, for a lower storage bill than the head sampler.
What the saving costs. The tail tier needs memory and compute. Using the earlier sizing, about 4.8 GB of buffer across Tier 2 plus headroom, say four replicas at 2 vCPU and 2 GB each, plus two small gateways. At an illustrative 30 to 50 USD per month per modest cloud vCPU-and-memory slice, the fleet might cost a few hundred USD per month. That is small against a saving of roughly 48,000 USD per month relative to storing everything, and about equal to the head-sampling bill in total. The network is the quieter cost: all 5.5 TB per day still crosses from applications to the gateway, and from gateway to Tier 2. In many clouds, cross-zone transfer is billed. If it is, place Tier 1 and Tier 2 in the same zone as the producers where your availability design allows, or measure the transfer charge before assuming the saving is clean.
The sensitivity that matters. The total kept fraction is dominated by the baseline rate and the error-or-slow share. If a bad deploy pushes the error rate from 1.5 percent to 20 percent, the tail sampler will happily keep 20 percent of traffic as error traces and your bill spikes at the exact moment you are least able to investigate it. This is why the rate-limiting and composite budget policies exist, and why the next section treats them as part of the design, not an afterthought.
Trade-offs, Gotchas, and What Goes Wrong
Tail sampling moves cost from storage to operations. The failure modes below are the ones that actually hurt.
Split traces. If routing by trace ID is bypassed, partial traces arrive at different samplers. Common causes: a second ingestion path that skips Tier 1, an agent configured to export straight to Tier 2, or a mismatch between the resolver’s view of the ring and reality during a rollout. Symptoms are policies that appear random: latency policies that never fire, because each sampler sees a fragment, or duplicate decisions on the same trace. Detect it by comparing the number of spans per decided trace with your expected distribution.
Rollout churn. Restarting Tier 2 discards buffered traces. Scaling it remaps trace IDs. Both cause a burst of incomplete decisions. Use rolling updates with a surge of capacity, enable the load balancer’s retry and queue settings, set a PodDisruptionBudget, and avoid autoscaling Tier 2 on short-term CPU signals. Some deployments scale it on memory utilisation of the trace buffer with a long stabilisation window.
Buffer overflow. When arrival rate times wait exceeds num_traces, the oldest traces are decided early. A traffic spike therefore degrades decision quality exactly when volume is highest. Alert on the sampling_trace_dropped_too_early metric and on memory approaching the limit set in the memory limiter. The memory limiter refuses data when the process is near its limit, which converts a hard OOM kill, and the total loss of the buffer, into backpressure. Put it first in the processor list.
Late spans. Spans that arrive after the decision are treated by the decision cache if enabled, and otherwise start a new buffer entry that is decided on its own, usually wrongly. The sampling_late_span_age histogram tells you how often this happens. Asynchronous workflows with minutes between spans are poorly suited to a 30-second wait. For these, consider a different identity for the unit of sampling or a longer-wait dedicated pipeline.
Cost spikes under failure. As computed above, an outage raises the kept fraction. Combine the always-keep error policy with a composite or rate-limiting guard, and decide in advance what you drop first when the budget is exceeded. A sensible order is to sample a proportion of identical errors rather than every one, because the thousandth identical stack trace adds little.
Instrumentation lies. The processor can only act on the status your spans carry. Services that swallow exceptions, or that return success codes for failed work, hide from the status_code policy. Fix instrumentation first; no sampler repairs missing signal.
Operational blind spots. The processor is documented as beta for traces, and several of its newer features are gated. Pin Collector versions, read the changelog before upgrades, and test config validation in CI. A misconfigured policy tends to fail quietly: the pipeline runs, but it keeps the wrong traces.
Latency added to visibility. Traces appear in the backend only after the wait, 30 seconds by default, plus batching. If on-call engineers expect traces seconds after a failure, tell them. Metrics and logs remain the fast signals; traces become the slower, richer follow-up.
Cross-signal correlation. Logs carrying a trace ID may point to a trace the sampler dropped. If your investigation workflow depends on “click from log to trace”, measure how often the link resolves. For GenAI and agent systems this is acute: long, nested traces with large payload attributes inflate buffer memory, a topic we examine in OpenTelemetry GenAI semantic conventions for LLM and agent observability. And if you complement application traces with kernel-level data, the trade-offs in eBPF observability and kernel tracing for APM change how much you need to sample, since eBPF capture has a different cost profile.
Practical Recommendations
Start by measuring before changing anything. Record your span rate, average spans per trace, bytes per span, error share, and the distribution of trace durations. Every sizing decision flows from those five numbers, and teams that skip this step choose a wait and a buffer by guess.
Next, build the data path before the policy. Deploy the two tiers, route by trace ID with the Kubernetes resolver, and run the sampler with a single always_sample policy for a week. This proves that routing works, shows real memory use, and reveals split traces, all with no loss of data. Only then replace the always-sample policy with a real set.
Order the policies by value. Errors first, then latency outliers per route, then business-critical paths at an elevated rate, then a small unbiased baseline, then a budget guard. Keep the baseline above zero; a sampler that stores only anomalies cannot tell you what normal is.
Move RED metrics ahead of the sampler on day one, so that dashboards and alerts never depend on stored traces. Treat sampling changes as production changes: roll out per environment, compare the kept fraction against prediction, and keep a rollback config.
- Measure span rate, trace size, and duration percentiles first.
- Run Tier 1 and Tier 2 with
always_samplefor a week to validate routing. - Set
num_tracesper replica from rate times wait, plus margin. - Alert on
sampling_trace_dropped_too_earlyand on memory limiter refusals. - Compute
span_metricsbefore sampling, with bounded dimensions. - Add a rate or composite guard so an outage cannot blow the budget.
- Pin Collector versions and validate configs in CI.
- Review the kept fraction and policy hit counts monthly.
Frequently Asked Questions
What is the difference between head sampling and tail sampling in OpenTelemetry?
Head sampling decides whether to keep a trace at the start, usually from the trace ID and a target percentage, so it cannot see errors or latency that happen later. Tail sampling waits until most spans of the trace have arrived and then decides using the full evidence, such as an error status or a long duration. Head sampling is cheap and stateless; tail sampling is more accurate for debugging but needs a stateful collector tier with enough memory.
Why must all spans of a trace reach the same collector for tail sampling?
The tail sampling processor evaluates policies on the spans it has buffered. If spans of one trace are split across instances, each instance sees a fragment and may judge it by incomplete evidence, dropping a failing trace because the error span went elsewhere. The load-balancing exporter solves this by routing on trace ID, so every span of a given trace lands on one back-end replica.
How do I size decision_wait and num_traces?
Choose decision_wait longer than the traces you want decided correctly, then estimate buffered traces as new traces per second multiplied by the wait. At 2,000 traces per second and a 30-second wait, about 60,000 traces are in flight across the tier. Divide by replicas for the per-instance num_traces, add margin, and then measure real memory per span, since span size and count drive the memory used.
Does tail sampling reduce the load on my applications?
No. Applications still create and export every span, because the decision happens later in the collector. Tail sampling reduces storage and backend ingestion cost, not instrumentation overhead or network traffic from the application to the first collector. If application overhead matters, combine it with a modest head sampling rate, remembering that anything the head sampler drops is invisible to the tail sampler.
Will sampling break my service metrics?
It will if you compute metrics from stored traces. Sampled traces under-count traffic and over-represent errors by design. Compute request, error, and duration metrics from all spans before the sampler, using the span metrics connector in a pipeline that taps the full stream, and send those metrics to your metrics backend. Then dashboards reflect true volume while traces remain a curated sample.
Can I use tail sampling and probabilistic sampling together?
Yes. Many deployments use a probabilistic policy inside the tail sampler for the healthy baseline, and some add a head or probabilistic stage earlier to protect the pipeline. For statistically correct counts across stages, the OpenTelemetry tracestate scheme with the th and rv fields lets samplers record the probability applied. That specification is still in development, so verify behaviour in your Collector version before relying on it.
Further Reading
- OpenTelemetry Collector architecture and pipelines for receivers, processors, exporters, and connectors.
- Grafana Alloy OpenTelemetry Collector tutorial for the same pipeline ideas in Alloy.
- OpenTelemetry GenAI semantic conventions for LLM and agent observability for large, deep agent traces.
- eBPF observability and kernel tracing for APM for a complementary capture path.
- OpenTelemetry sampling concepts for the official definitions of head and tail sampling.
- Tail sampling processor README for the current policy list and settings.
- Probability sampling with tracestate for the th and rv specification.
By Riju — about
