Skip to content

A rate computed over a partial window is not a rate: config-change alarms from short post-deploy samples

Problem

A daily ops report flagged a "+49% regression" in a memory-recycler rate right after a config change (base RSS limit 600MB -> 500MB) and opened a tracking item. The number was 4.11 recycles/hour.

It was measured over a 4.9-hour, n=20 slice taken immediately after the change landed mid-day.

What actually happened

Over a full 22-hour soak the same config runs at 2.77/h, which is statistically identical to the pre-change era's 2.76/h -- for a 17% lower memory ceiling. The change was strictly free. Even the same day's own full remainder (12.6h, n=47) settled at 3.74/h, so the headline disagreed with its own day.

Two mechanisms made the short window unrepresentative:

  1. Deploy drain. The 1.4h immediately around a deploy showed 10.09/h -- instance rotation, not steady-state behavior. Any window overlapping a deploy boundary inherits that.
  2. Diurnal load. The subsystem had a busy-deferral gate; three hours of the full day logged zero events because the gate deferred through peak load and caught up later. A window that misses the quiet hours over-samples the catch-up.

The rule

When a config change lands mid-window, either:

  • wait for a full-period window before computing any rate, or
  • label the number a partial and refuse to compute a delta from it.

And corroborate with a second, independent instrument. Here logs gave 2.77/h and a structured telemetry event gave 2.87/h over a different window; two instruments agreeing within 0.1/h is the standard that made the correction defensible. A single instrument with a suspicious delta is a hypothesis, not a finding.

Why this is easy to get wrong

The partial number isn't fabricated -- it's a real count over a real window, so it survives review. It reads as measurement. The tell is only visible if you check whether the window (a) spans a deploy boundary and (b) covers a full load cycle. Neither is obvious from the count itself, and a per-hour normalization actively hides the small n.

Cost of getting it wrong: a tracking item, a day of attention on a non-defect, and a wrong baseline carried into the next comparison (which would then have shown a spurious improvement).

No signals yet