Skip to content

Sentry spans dataset: count() is extrapolated by 1/sampling_rate, count_sample() is the raw count - and the two differ 31x on a sampled project

TL;DR.

In Sentry's EAP spans dataset, count() silently scales each span by 1/client sample rate; count_sample() is the unscaled count. Measured 31x inflation on a project at tracesSampleRate 0.03, exact parity at 1.0. Extrapolation also cannot see before_send_transaction filtering, so it does not compose with duration-ladder sampling.

Symptom: a Sentry dataset=spans query returns a count() that is wildly higher than the transactions you know you sent, with no indication anything was scaled. Measured on one of our low-traffic frontend projects: count() 5,913 vs 189 actually ingested over 7d - a 31x inflation that reads as real traffic.

Cause: count() on the spans dataset (EAP) is an EXTRAPOLATED metric. Sentry scales each span by 1 / sampling_rate, where sampling_rate is the client-reported rate from the trace context (the DSC), so a project running tracesSampleRate: 0.03 reports 33x its measured volume by design. This is intentional and documented in passing in the transactions->spans migration FAQ ("the spans dataset has extrapolation built in to account for your client side sample rates"), but nothing in the API response or the UI number marks the value as estimated.

The escape hatches, none of which are obvious:

  • count_sample() - the RAW, unscaled count. This is what you want for any parity check, quota reasoning, or de-biasing math.
  • sampling_rate - queryable as a field: avg(sampling_rate), min(sampling_rate). Use it to detect whether extrapolation is even active before trusting count().
  • Filtering on it needs a numeric literal; sampling_rate:<1 is rejected with "<1 is not a valid filter value for sampling_rate, expecting double, but got a <class 'str'>".

Diagnostic recipe - run both counts side by side, per environment, and compare:

GET /api/0/organizations/{org}/events/
  ?dataset=spans&statsPeriod=7d&project={id}&environment=prod
  &query=is_transaction:1
  &field=count()&field=count_sample()&field=avg(sampling_rate)&field=min(sampling_rate)

Real output across three projects in one org, same query:

prod   count()=3,958   count_sample()=3,958   avg=1.0    min=1.0    -> no extrapolation
stage  count()=10,864  count_sample()=10,864  avg=1.0    min=1.0    -> no extrapolation
dev    count()=5,279   count_sample()=5,198   avg=0.985  min=0.1    -> slight, from one low-rate client
gt-web count()=5,913   count_sample()=189     avg=0.032  min=0.01   -> 31x

Three traps this creates:

  1. An unscoped query silently mixes extrapolation regimes. Omit the environment filter and a handful of local dev requests at rate 0.1 inflate an otherwise-exact production number. Always scope by environment, or use count_sample().
  2. Extrapolation cannot see post-hoc filtering. If you head-sample at 1.0 and then discard 99% of fast transactions inside before_send_transaction (the standard duration-ladder pattern), the client still reports sampling_rate: 1.0, so Sentry applies NO correction and the UI undercounts real traffic by up to 1000x. Extrapolation and a duration ladder do not compose: whichever one you rely on, the other is invisible to it. Your own inverse-keep-rate weights remain necessary, and they must be applied to count_sample(), not to count(), or you double-correct.
  3. It changes meaning silently across projects in the same org. Two projects, same query, one exact and one 31x - so a shared dashboard or a cross-project script is comparing incomparable numbers.

Rule of thumb: in the spans dataset, treat count() as an estimate and count_sample() as the measurement. Any script that feeds Sentry counts into capacity, cost, or regression math should use count_sample() and assert on min(sampling_rate) so a future sampling change fails loudly instead of quietly rescaling the output.

No signals yet