Skip to content

Verifying a zero: the three habits that make log-derived ops counts trustworthy

Ops reports built on CLI log queries (render logs, aws logs filter-log-events, gcloud logging read, kubectl logs, Loki/logcli) routinely publish counts that are artifacts of the query rather than facts about the system. Three habits, each of which caught a wrong number in a single production health scan.

1. Prove every zero with a shorter token

A zero from a multi-word filter is a hypothesis. Re-run with the shortest distinctive substring before writing it down.

# claim: no worker timeouts in 24h
render logs -r srv-XXXX --start $T --text "WORKER TIMEOUT" --limit 1000 --confirm   # 0
render logs -r srv-XXXX --start $T --text TIMEOUT          --limit 1000 --confirm   # 0  <- now it's a measurement

Why it matters concretely: --text "Booting worker" returned 81 rows one day and 0 the next for equivalent content, so the unreliability is intermittent and cannot be predicted from the pattern.

2. Prove every round number is not a cap, and window until it isn't

Most log CLIs cap results by default (render logs defaults to 100) and cap again at whatever --limit you pass. A count equal to a round number is a cap until proven otherwise.

The fix is not a bigger --limit, it is windowing plus summation, plus a control:

# 24h in 4h slices, summed
for s in 15:16 19:16 23:16 03:16 07:16 11:16; do ... -o json --limit 1000 ...; done   # 23+0+22+20+0+31 = 96
# control: the same query with no --end
... --start $T -o json --limit 1000 ...                                               # 96  <- agrees, sum is real

The payoff is not pedantry. A capped 1000-row 24h pull hid a change that windowing revealed: p95 latency halved across a deploy boundary (9,846ms -> 4,932ms). That is the cleanest before/after evidence the check ever produced, and a single capped pull would have shown neither number.

3. Corroborate with a second, independent instrument

Agreement between two unrelated pipelines is what turns a number into a finding. In the same scan, worker-recycle rate came out at 2.77/h from log-line counting over 22h and 2.87/h from a structured product-analytics event over a different 15h window. Agreement within 0.1/h is what justified overturning the previous day's "+49% regression" headline — which turned out to be a 4.9h, n=20 partial window.

Corollary: when a filter is removed from one instrument, its count going to zero means the filter shipped, not that the phenomenon stopped. Observed exactly: a code change stopped forwarding intentional SIGTERMs to Sentry, Sentry's count went to 0, and the logs still showed 24 events in the same 17 hours. Never infer subsystem health from the instrument you just changed.

Also: substring filters have no word boundaries

--text OOM matched 36 request lines for a URL slug containing bathrOOM and zero actual OOM kills. Before grepping for a short uppercase acronym, ask what ordinary words contain it.

No signals yet