Doing daily product-analytics KPI reads on a low-volume site (PostHog, but the shape is provider-agnostic), the standard bot filter is a screen-resolution denylist such as NOT (screen 800x600). Two failure modes bit us on 2026-08-15.
1. A run of zero bot days is not evidence the filter is unnecessary. The filter returned zero rows for five consecutive days, then returned 9 pageviews on day six. Those 9 were 19% of the day's total pageviews and 9 of 12 home-page views, so the home path's apparent 5 -> 12 jump was entirely synthetic. Uncorrected, the day's external KPI over-counted by 19% (47 -> 38 pageviews, 34 -> 25 unique visitors) and would have been reported as +1.12x when it was actually 0.90x against the trailing-7d median. At low volume a single small fleet flips the sign of the headline number. Always run the filter and always report the raw AND corrected figures when bot rows are nonzero.
2. The strongest single tell is viewport LARGER than screen. The fleet reported screen 800x600 with viewport 1600x1200. A real browser cannot render a viewport bigger than the physical screen; headless Chrome routinely reports a default screen size while the page is rendered at whatever the automation set. This one comparison is a cheaper and more robust discriminator than any resolution denylist, because it needs no list to maintain and catches fleets that pick a plausible screen size.
The corroborating fingerprint, all present together:
- every hit on one path (here
/), all$directreferrer - one distinct_id per pageview (no session has two events)
- sequential distinct_id blocks (
01a000b2,b5,b8,bb,be,c2,c8,cc) -- ids allocated from a counter, not randomly - a burst at fixed spacing (8 of 9 arrived 3-4 minutes apart across 28 minutes)
- desktop Chrome/Edge on Windows, i.e. the default UA of most automation stacks
One id per pageview plus sequential ids is decisive on its own: real visitors reuse an id across a session, and real id allocation is random. This shape is an uptime monitor or crawler fleet, not people.
Scope caveat worth stating in any report built on this data: a JS-side analytics tool only sees clients that execute JS, so a non-JS bot flood remains invisible no matter how good the filter is. Cross-check server-side request logs or bandwidth before claiming a clean anomaly section.