Symptom: a Sentry dataset=spans query returns a count() that is wildly higher than the transactions you know you sent, with no indication anything was scaled. Measured on one of our low-traffic frontend projects: count() 5,913 vs 189 actually ingested over 7d - a 31x inflation that reads as real traffic.
Cause: count() on the spans dataset (EAP) is an EXTRAPOLATED metric. Sentry scales each span by 1 / sampling_rate, where sampling_rate is the client-reported rate from the trace context (the DSC), so a project running tracesSampleRate: 0.03 reports 33x its measured volume by design. This is intentional and documented in passing in the transactions->spans migration FAQ ("the spans dataset has extrapolation built in to account for your client side sample rates"), but nothing in the API response or the UI number marks the value as estimated.
The escape hatches, none of which are obvious:
count_sample()- the RAW, unscaled count. This is what you want for any parity check, quota reasoning, or de-biasing math.sampling_rate- queryable as a field:avg(sampling_rate),min(sampling_rate). Use it to detect whether extrapolation is even active before trustingcount().- Filtering on it needs a numeric literal;
sampling_rate:<1is rejected with"<1 is not a valid filter value for sampling_rate, expecting double, but got a <class 'str'>".
Diagnostic recipe - run both counts side by side, per environment, and compare:
GET /api/0/organizations/{org}/events/
?dataset=spans&statsPeriod=7d&project={id}&environment=prod
&query=is_transaction:1
&field=count()&field=count_sample()&field=avg(sampling_rate)&field=min(sampling_rate)Real output across three projects in one org, same query:
prod count()=3,958 count_sample()=3,958 avg=1.0 min=1.0 -> no extrapolation
stage count()=10,864 count_sample()=10,864 avg=1.0 min=1.0 -> no extrapolation
dev count()=5,279 count_sample()=5,198 avg=0.985 min=0.1 -> slight, from one low-rate client
gt-web count()=5,913 count_sample()=189 avg=0.032 min=0.01 -> 31xThree traps this creates:
- An unscoped query silently mixes extrapolation regimes. Omit the environment filter and a handful of local dev requests at rate 0.1 inflate an otherwise-exact production number. Always scope by environment, or use
count_sample(). - Extrapolation cannot see post-hoc filtering. If you head-sample at 1.0 and then discard 99% of fast transactions inside
before_send_transaction(the standard duration-ladder pattern), the client still reportssampling_rate: 1.0, so Sentry applies NO correction and the UI undercounts real traffic by up to 1000x. Extrapolation and a duration ladder do not compose: whichever one you rely on, the other is invisible to it. Your own inverse-keep-rate weights remain necessary, and they must be applied tocount_sample(), not tocount(), or you double-correct. - It changes meaning silently across projects in the same org. Two projects, same query, one exact and one 31x - so a shared dashboard or a cross-project script is comparing incomparable numbers.
Rule of thumb: in the spans dataset, treat count() as an estimate and count_sample() as the measurement. Any script that feeds Sentry counts into capacity, cost, or regression math should use count_sample() and assert on min(sampling_rate) so a future sampling change fails loudly instead of quietly rescaling the output.