Adaptivetailsampling Processor
Supported Telemetry
Overview
The Adaptive Tail Sampling Processor performs adaptive tail-based trace sampling using first-match rules-based routing to adaptive samplers. Each sampler produces a known sample rate which is encoded asot=th in W3C TraceState for correct downstream metric weighting.
How it works
- Spans are accumulated in memory, grouped by trace ID.
- A trace becomes ready for evaluation when any of three triggers fires (whichever comes first):
- a root span arrives (any span with an empty
ParentSpanID), trace_timeoutelapses since the first span of the trace arrived. This timer is set on first-seen and is never extended by subsequent spans, ensuring a predictable upper bound on buffer occupancy, or- the trace accumulates
span_limitspans (10000 by default; 0 disables). This bounds the memory a single giant trace can hold; the decision is made over the spans buffered so far, and spans arriving afterwards are stamped from the decision cache instead of being buffered.
- a root span arrives (any span with an empty
- After the trigger fires, the processor pauses for
decision_delayto let in-flight straggler spans land. The root-span and trace-timeout triggers share the same delay; the span-limit trigger decides immediately, since the trace is at its memory peak and waiting would let it keep growing. - Rules are then evaluated in order against the accumulated trace. The first rule whose conditions all match selects the sampler; once a rule is selected, its sampler’s keep/drop decision is final and no later rules are considered. A rule with no conditions is a catch-all.
- The matched sampler produces a sample rate (1-in-N). Samplers only ever produce rates; no sampler makes a keep/drop decision itself.
- The processor converts the rate to a threshold (a rate of N becomes the threshold encoding probability 1/N) and makes the keep/drop decision by comparing that threshold against the trace’s randomness (
ot=rvwhen present, otherwise derived from the trace ID), using the OTel consistent probability sampling algorithm.
[!NOTE] The adaptive samplers come from dynsampler-go, which the processor uses only to compute rates. dynsampler-go’s own samplers can make keep/drop decisions internally (itsDeterministicSamplerapplies its own hash-based check, for example), but this processor never uses that path: the rate-to-threshold conversion and the randomness comparison in step 6 are the single decision mechanism for every sampler type, including this processor’sprobabilistictype. This is what makes decisions reproducible for a given trace and correctly weighted downstream viaot=th.
- Sampled traces are forwarded with annotations on every span:
otelcol.processor.adaptive_tail_sampling.rule: the name of the matched ruleotelcol.processor.adaptive_tail_sampling.trigger: which event triggered the decision (see Output attributes)- W3C TraceState
ot=th:<hex>: the threshold encoding the effective sample rate
- The decision (sampled or dropped) is recorded in a per-trace LRU cache so that any late-arriving spans for the same trace ID are handled consistently with the original decision (see Decision cache below).
Configuration
Root-span detection
The processor moves a trace out of accumulation and into the decision-delay phase as soon as it sees a span it considers a “root”. By default that meansIsRootSpan(): any span whose W3C ParentSpanID is empty. This works for producer-side head-sampled traces where the true root always lands at the same collector, but it stalls when:
- Only the server side of a cross-process trace reaches this collector, so no observed span has an empty parent.
- The operator wants to trigger on a producer-supplied hint (e.g. a message-broker consumer or batch-job entry point) rather than on the transport-level root.
root_span_condition lets the operator override the trigger with any OTTL boolean expression evaluated in the ottlspan context. The first span for which it returns true starts the decision-delay timer.
Producer hint (fires on any real root, or on any span the producer explicitly tagged):
trace_timeout):
processor_adaptive_tail_sampling_ottl_eval_errors under the sentinel rule="_root_span_condition" label so they show up separately from rule-condition errors. A span whose evaluation errored is treated as non-matching, so a broken expression can only delay the trigger to trace_timeout, never fire it prematurely.
Rules
Rules are evaluated in order; the first whose conditions all match selects the sampler. A rule with no conditions is a catch-all. Once a rule is selected, its sampler’s decision is final: a drop stays dropped, the trace is not handed to any later rule. Each rule’sname must be unique. It is recorded on the decision metrics (the rule attribute) and stamped on sampled spans, so pick names you want to see on dashboards. Names starting with _ are rejected at config validation: that prefix is reserved for processor-internal decision labels (such as _eviction), so user rules can never collide with them.
[!WARNING] A rule with no conditions (a catch-all) placed before another rule consumes every trace and renders the later rules unreachable. The processor logs a warning at startup when it detects this configuration so it shows up in collector logs.Conditions are OTTL boolean expressions evaluated in the
ottlspan context, which gives access to the span, its resource, and its instrumentation scope. Path expressions should be qualified with the context they refer to (span.attributes["k"], resource.attributes["k"], span.status.code, and so on); unqualified paths are valid OTTL and resolve against the span context. We recommend qualifying every path, all examples in this README do, and it keeps conditions visually consistent with fingerprint selectors, which require a scope. See OTTL Boolean Expressions for the full grammar (including and, or, and parentheses), and the ottlspan paths reference for the list of accessible fields.
The intended pattern is specific-conditions rules first, catch-all last:
adaptive_percentage. Flipping the order so default comes first would mean default swallows every trace (including errors) and keep-errors is never reached, which is what the startup warning flags.
How multiple conditions combine
Each rule can carry multiple OTTL conditions. Thematch field controls how they are combined against the spans of the accumulated trace:
match: any_span(default): each condition must be satisfied by some span in the trace, but not necessarily the same span across conditions. Reads as “the trace has these characteristics.”match: same_span: some single span in the trace must satisfy all conditions at once. Reads as “there is a span with these characteristics.”
match is not permitted on catch-all rules (rules with no conditions).
Example: keep a trace only when a single span is both an error and from the payment service:
any_span): keep any trace that touched the payment service and saw an error somewhere:
otelcol_processor_adaptive_tail_sampling_ottl_eval_errors counter is incremented, labelled by the rule name. Other spans and conditions continue to evaluate normally.
Condition examples
Keep every trace where an HTTP server span returned 5xx, using OTTL’s inlineand:
IsRootSpan() helper:
IsMatch regex helper:
Samplers
Each type names a distinct intent:
Decisions are always made per trace; volume is measured in spans, so a kept
trace counts all of its spans against a throughput budget, and a percentage
goal is measured over span volume (equal to the percentage of traces when
trace sizes are uniform).
The
adaptive_* samplers are backed by
dynsampler-go and accept an
optional algorithm selecting how per-key rates are computed: ema (the
default) smooths traffic with an exponential moving average, tuned with
adjustment_interval and weight; windowed (supported by
adaptive_throughput only) computes rates over a sliding window, reacting
faster to traffic shifts at the cost of more spike sensitivity, tuned with
update_frequency and lookback_frequency.
Migrating from Refinery: DeterministicSampler -> probabilistic (the same
hash-consistent fixed fraction), EMADynamicSampler -> adaptive_percentage
(note Refinery’s GoalSampleRate: N means keep 1-in-N, so GoalSampleRate: 5
becomes goal_percentage: 20), EMAThroughputSampler -> adaptive_throughput,
WindowedThroughputSampler -> adaptive_throughput with algorithm: windowed.
probabilistic
The inline equivalent of the probabilistic_sampler processor.
adaptive_percentage
Adapts the sample rate per traffic key over time, keeping a target average
percentage across all keys: rare keys survive, chatty keys are sampled
aggressively.
adaptive_throughput
Adjusts rates per key to hit a sustained volume budget in spans per second.
[!IMPORTANT]
goal_throughput is enforced per collector instance. Each instance targets the goal against the traffic it sees, so a fleet of N instances emits up to N times the configured throughput. Divide the backend budget by the instance count when sizing this value. See Deployment considerations.
With algorithm: windowed, rate recalculation (update_frequency) is
decoupled from the historical window used for the calculation
(lookback_frequency):
Fingerprints
Thefingerprint_attributes field names the attributes that identify what kind of trace this is for sampling purposes. Each distinct fingerprint value gets its own adaptive sample rate, so choose attributes that classify traffic (route, status code, method, service) rather than identify individual requests (user IDs, request IDs, raw URLs), which would give every trace its own key and defeat the adaptation.
Each entry is a scoped attribute selector of the form <scope>.attributes["<name>"]:
The
resource., scope., and span. prefixes match OTTL’s span-context path names, so conditions and fingerprint entries share one spelling; root. and any. are trace-level scopes OTTL cannot express (fingerprints are built from the whole trace, while OTTL evaluates one span at a time). Note that fingerprint entries are selectors, not OTTL expressions. The two concepts have distinct jobs throughout the config: OTTL conditions appear wherever a single span is evaluated (conditions, root_span_condition), and selectors appear wherever a trace-level value is collected. Qualifying every path in conditions keeps the two styles identical in practice.
any. is simply the widest search, and a value present at several origins appears once. The fingerprint for a trace is built by sorting the distinct values each selector matched and joining them with , within each entry, then joining the entries with the • separator. A trace whose spans carry several values for one selector keys as the combination (e.g. checkout,billing•/api), which is worth knowing when debugging unexpectedly high key cardinality. Selectors that match nothing are replaced with <missing>.
Extraction cost scales with the scope: resource. and scope. are independent of trace size, while span., any., and root. walk every span of the trace at decision time. root. additionally evaluates the root-span condition per span, which is cheap for the default condition but costs an OTTL evaluation per span for custom ones.
Worked examples
Two complete configurations for the most common deployment shapes.Error retention with an adaptive default
Keep every error trace, and let an adaptive sampler settle the rest at a target percentage per traffic class. This is the right starting shape for most single-instance deployments: errors are never lost, and the default rule adapts as traffic mix shifts.num_traces bounds memory and should cover the number of traces that start within one trace_timeout window at peak. The decision caches should be a small multiple of that, since they only store trace IDs and outcomes; undersizing them turns late spans of decided traces back into new partial traces. num_traces bounds how many traces are buffered, not how large any one of them grows; span_limit (10000 by default) covers that dimension. Keep it well above your largest legitimate trace so a single runaway trace cannot exhaust memory while normal traces are never truncated.
Throughput-bounded fleet
Cap the spans-per-second this collector emits regardless of incoming volume, useful when the downstream backend is provisioned for a fixed ingest rate. The throughput sampler continuously adjusts per-key rates to hold the total near the goal.service.name means each service’s share adapts to its share of total traffic rather than being fixed.
Decision cache
When a trace’s decision is finalised, the trace ID and outcome are recorded in one of two LRU caches. Any later spans that arrive for that trace ID short-circuit the accumulation path and are handled consistently with the original decision:- Sampled cache (
decision_cache.sampled_cache_size): late spans are forwarded immediately, stamped with the same rule attribute andot=thTraceState as the original batch. - Not-sampled cache (
decision_cache.non_sampled_cache_size): late spans are silently dropped.
0 disables that side of the cache. Late spans for a
decision class whose cache is disabled fall through to the normal pending-trace
path and may produce a second (possibly inconsistent) decision; with both caches
disabled, every late span falls through. Operators tune the cache sizes against
observed late-span volume.
Buffer overflow and eviction policy
The in-memory accumulation buffer is sized bynum_traces. When the buffer is
full and spans for brand-new traces arrive, the processor evicts the oldest
pending traces to make room (the buffer may transiently exceed num_traces by
the number of new traces in a single incoming batch). Every evicted trace
receives a real sampling decision immediately, with the spans seen so far and
no decision_delay; the decision is recorded in the decision cache so
late-arriving spans are handled consistently, and kept traces carry correct
ot=th annotations. How that decision is made is configurable:
evaluate(default): the evicted trace runs through the normal rules and sampler path. Your rules keep working under pressure (e.g. a keep-errors rule still keeps error traces), at the cost of OTTL evaluation over every span of every evicted trace.probabilistic: rule evaluation is skipped and the trace is decided by comparing a threshold derived fromsampling_percentageagainst the trace’s randomness. Constant-time work per evicted trace regardless of trace size or rule count, so an overloaded instance sheds load instead of amplifying it. Kept traces are stamped with the correspondingot=th(and the sentinel rule attribution_eviction), so downstream weighting stays accurate. Recommended for high-throughput deployments where eviction indicates genuine overload.
decision_cache.*) is a separate
structure recording completed decisions; it does not protect pending traces
from eviction.
Operators sizing the processor should watch traces_evicted (every eviction)
and decision_triggers{trigger="eviction"} and increase num_traces if they
are non-zero in steady state.
Deployment considerations
The processor accumulates spans in memory, so all spans of a given trace must reach the same processor instance. Multi-instance deployments use the same two-tier pattern as thetail_sampling processor:
adaptive_throughput (with either algorithm) this means goal_throughput is a per-instance target: a fleet of N instances emits up to N times the configured goal, so divide the backend’s ingest budget by the instance count when sizing it. The percentage-based samplers (adaptive_percentage, probabilistic) are unaffected, since a target percentage composes across instances. Automatic cluster-size awareness for the throughput goal is not currently implemented.
Known limitations
- Decisions are per-instance. All spans of a trace must reach the same
instance; scale out with the two-tier
loadbalancingpattern above. Shared storage for a single-tier deployment is future work. - The buffer is bounded by trace count, not bytes. Steady-state memory is
approximately span rate x buffer window (
decision_delayafter the decision triggers, up totrace_timeout) x average span size, plus the decision caches. Trace shape barely matters: per-trace bookkeeping is small, so many small traces cost slightly more than a few giant ones at the same span rate. Whennum_tracesis exceeded the oldest trace is evicted with a real decision (see the eviction policy section). - Shutdown drains rather than discards. Pending traces are decided with the spans seen so far and kept traces are forwarded, so a clean restart or rollout does not silently lose the buffer. A crash still loses it.
- Rules are fixed at startup. Changing rules means a restart (which drains, see above). Hot reload is future work.
- Traces only. Rule conditions use the OTTL span context, and the configuration is trace-scoped; correlated log sampling is future work and will add log-specific configuration alongside the existing fields.
Relationship to processor/tail_sampling
This processor sits next to tail_sampling rather than extending it. The two are close in shape (both buffer traces and decide once a trigger fires) but use different evaluation models, and tail_sampling’s current model has several mechanical features that a rate-bearing sampler does not fit cleanly into.
Decision semantics
Intail_sampling’s policy loop only two outcomes short-circuit evaluation: any policy returning Dropped, and (when sample_on_first_match is enabled) the first policy to return Sampled. Every other outcome, including NotSampled, lets the loop continue to subsequent policies.
All common “this trace did not pass my check” votes from existing policies (probabilistic, status_code, rate_limiting, and_policy, etc.) return NotSampled, never Dropped. Dropped is reserved for the explicit drop policy, and the final-decision composition gives Dropped precedence over Sampled.
A rate-bearing sampler dropped into this model has to pick a return value for a probabilistically-dropped trace, and both options break correctness:
- Returning
NotSampledmatches the existing convention but does not stop the loop. A later policy votingSampledwould still cause the trace to be kept, and the rate-bearing sampler’s view of what it controlled diverges from reality. Its rate calculations drift over time. - Returning
Droppedstops the loop, but it also wins precedence over everySampledvote in the composition step. An operator pairing a rate-bearing sampler with an explicitkeep-errorspolicy would expect errors to always win; instead the rate-bearing sampler’s probabilistic drop would override the keep, inverting the configured intent.
No rate or threshold in the policy contract
TheEvaluator interface returns only a decision enum. There is no way for a policy to communicate the sample rate or threshold it applied, which means ot=th cannot be emitted from inside tail_sampling without expanding the contract. The processor has roughly twenty existing policies; widening the interface would touch all of them.
sample_on_first_match would become correctness-load-bearing for one policy type only
Today sample_on_first_match is an opt-in optimization. Adding rate-bearing samplers as policies would make it mandatory for configurations using them; without it, the OR composition above produces drift. Existing tail_sampling users adding a rate-bearing policy without flipping the flag would silently get wrong rates. Validation cannot easily reject this because the flag is fine in isolation, and there is no current marker on a policy type that says “this policy requires first-match for correctness.”
Rule attribution
This processor records the matched rule name on every span in a sampled trace, which is only meaningful under first-match semantics. Undertail_sampling’s multi-policy OR model multiple policies can vote Sampled and there is no defined “winner” to attribute the trace to.
Relationship to PR #48865
In-flight work on adding tracestate handling totail_sampling’s probabilistic policy is input-driven: it reads SDK-supplied ot=th to adjust the tail probabilistic threshold (equalizing mode). This processor is output-driven: it computes a tail-stage rate and emits ot=th. The two address different problems and can ship independently.
For these reasons adaptive_tail_sampling is a separate processor. The tail_sampling users retain the existing multi-policy model unchanged, and adaptive_tail_sampling keeps a single evaluation model (first-match with rate-bearing samplers) end to end.
Interoperability with upstream sampling
The processor honours any incomingot=th (sampling threshold) and ot=rv (explicit randomness) already set on incoming spans by an upstream sampler, such as an SDK head sampler or a probabilistic collector processor:
- Randomness. If an incoming span carries
ot=rv, that value is used to make the sampling decision. Otherwise the last 7 bytes of the trace ID are used, per the consistent probability spec. - Population-relative rate (equalizing). The rule’s rate
Nis interpreted as the operator’s target for the original population: “keep 1-in-N of all traces before any sampling.” The effective absolute keep probability ismin(P_upstream, 1/N). This is the same composition mode asequalizinginprocessor/probabilisticsamplerprocessor. - Threshold monotonicity. If a span already carries an
ot=thstricter than what this processor would emit, the incoming value is preserved. This matches the consistent probability spec: a downstream stage may raise a threshold but never lower it.
Worked example: composing with a head sampler
ot=th always reflects the
effective end-to-end probability:
- With the 10% goal above, roughly 10% of the original traffic survives (not 10% of the upstream sampler’s 50%), and every kept span carries the 10% threshold, so adjusted counts reconstruct the original volume.
- If the rule were looser than upstream (e.g.
always_sample), all arriving spans are kept and retain the upstream 50% threshold, so adjusted counts remain honest.
Accuracy under non-uniform upstream sampling
The equalizing composition above is exact when upstream sampling is uniform across the keys the rule’s adaptive sampler uses (fingerprint_attributes). If upstream head-samples different classes of traffic at different rates and those classes overlap with the tail sampler’s keys, the adaptive samplers observe a population that under-represents heavily-downsampled keys and can misjudge their per-key rate. Improving accuracy in that case requires per-key upstream tracking in the sampler and is tracked as follow-up work in #49517. Under uniform upstream sampling (the common case, e.g. an SDK TraceIdRatioBased sampler) the rates are exact.
Grouping unrelated traces with a shared ot=rv
Because the sampling decision is deterministic against the 56-bit randomness value, an upstream producer that sets the same ot=rv on multiple otherwise-unrelated traces will get the same sampling decision for all of them at the same threshold. This is useful when a set of traces should be sampled together as a group, for example:
- An LLM agent producing multiple traces for turns in a single conversation, all stamped with a
conversation.id-derivedot=rv. - A browser SDK producing multiple traces during one user session, all stamped with a
session.id-derivedot=rv. - A batch job producing one trace per task, all stamped with the batch’s job ID.
ot=rv value is a 56-bit number, so a stable hash of the entity ID truncated to 56 bits is a suitable derivation. The processor itself does not compute ot=rv from arbitrary attributes: the producer or an earlier processor is expected to set it. Whatever ot=rv is present when the accumulated trace arrives will be used for the decision and preserved on the emitted spans.
Multi-instance deployment considerations
Under the standard two-tier deployment pattern (loadbalancing exporter routing traces to a downstream tier running this processor), how traces are routed interacts with rv-based grouping in two useful ways:
- Routing by trace ID (default). Traces in the same rv-group land on different collector instances because their trace IDs are unrelated. Sampling decisions remain correct without any coordination between collectors: because our decision is deterministic against the shared rv, every instance reaches the same keep/drop outcome for a given rv. The trade-off is that the adaptive samplers (
adaptive_percentage,adaptive_throughput) each see only a slice of the group’s traffic, so per-key rate calculations converge more slowly than if the whole group were visible to one instance. - Routing by a group-carrying attribute. If the producer sets both a shared
conversation.id(or similar) and the derivedot=rv,loadbalancingcan be configured withrouting_key: attributesnaming that attribute so all traces for one group land on the same collector. Adaptive-sampler learning is coherent across the group at the cost of load distribution: heavy-tailed group sizes (one active conversation among many quiet ones) skew load onto specific instances, andnum_traceson the busy instance must be sized for the largest concurrently-pending group ortraces_evictedwill start climbing. Group size is upstream-controlled, so an unexpected traffic spike within one group shifts the skew unpredictably.
num_traces can be sized conservatively.
Future work on shared trace context across collector instances (tracked under “Cross-instance shared state” in #49311) would remove this trade-off: with a shared backing store the adaptive samplers can learn from the whole group regardless of which instance decides any individual trace, so uniform-load routing (route by trace ID) no longer costs sampler learning coherence.
Metrics
The
rule label carries the matched rule’s name from the config. Values
prefixed with _ are processor-owned sentinels rather than user rules:
_unmatched (dropped with no matching rule), _eviction (decided by the
probabilistic eviction policy), and _root_span_condition (on
ottl_eval_errors, when the root-span condition itself fails to evaluate).
Config validation rejects user rule names starting with _, so the two can
never collide. These labels and the trigger values above are part of the
component’s telemetry contract; build dashboards on them freely.
Output attributes
Every span in a sampled trace is annotated with:Recording the fingerprint
record_fingerprint (default none) stamps the matched rule’s fingerprint on every span of a kept trace, the same way the rule name is recorded:
valuerecords the raw fingerprint (egcheckout,billing•/api). Long combination keys grow the sampled decision cache, which stores the recorded string for late spans.hashrecords the first 8 bytes of the fingerprint’s SHA-256 as 16 hex characters. The size is fixed and the hash is deterministic across instances and restarts, so grouping works fleet-wide. Verify a span’s hash by recomputing it from the raw fingerprint,echo -n '<fingerprint>' | sha256sum | cut -c1-16.
fingerprint_attributes (always_sample, probabilistic) never produce the attribute, and enabling either mode adds one attribute write per span on kept traces.
The sample rate is encoded in W3C TraceState as ot=th:<hex> per the OTel consistent probability sampling spec. The spanmetrics connector (enable_metrics_sampling_method: true) reads this field to produce correctly weighted R.E.D metrics from sampled data.
Future work
- Shared-storage backed scaling for single-tier deployments
- Rule and sampler hot-reload
- Span-count decision trigger (SpanLimit parity)
- Stress-relief style overload activation
- Correlated log sampling
Last generated: 2026-08-31