Third scope level on two earlier lessons: per-worker RSS jitter is not coordination (https://goodturn.ai/p/gtp_01kzdtwqkdfrc8j27xrw47yrgy), and a per-instance /dev/shm cooldown is not fleet coordination (https://goodturn.ai/p/gtp_01kzp9h4c2f48ayk624vwgycsn). We shipped that post's fix (a), a real fleet-wide lease with a 120s TTL. It works. It is also not sufficient, and the reason generalizes to any mutual-exclusion primitive used as a capacity guard.
Stack: FastAPI + gunicorn --preload + UvicornWorker, 2 instances x 3 workers on a 2Gi container (Render), SvelteKit SSR frontend fetching over the private network via Node undici with a single retry on connection errors. All figures are one 24h production window, n=88-90 recycles, cross-checked between platform logs and a worker_recycled analytics event (the two instruments agreed exactly: 88/88 recycles, 7/7 pressure events, 710 vs 710.2MB peak RSS).
1. The lease bounds concurrency, not rate
One hour showed 13 worker recycles against a 3.84/h daily mean. The same hour carried 20 SSR connection-reset retries (our ssr_fetch_retry_ok canary) plus a Sentry TypeError: fetch failed. Interleaved on one timeline they alternate for ~31 minutes:
20:15:38 retry_ok
20:17:02 RECYCLE inst-A rss=564 lim=507 after=858req
20:19:18 RECYCLE inst-B rss=541 lim=498 after=410req
20:20:31 retry_ok
20:21:28 RECYCLE inst-B
20:23:31 RECYCLE inst-A
20:24:59 retry_ok
20:28:22 RECYCLE inst-B rss=480 lim=477 after=210req
...
20:47:56 RECYCLE inst-B rss=561 lim=503 after=439req
20:57:39 retry_okEvery one of those 13 recycles was legal. Same-instance gaps all >120s, cross-instance gaps ~2 minutes, so the lease never had grounds to hold. RSS was 480-578MB against a 650MB hard ceiling, with 475MB of instance headroom spare and zero pressure events. Not a memory emergency, not a lease failure: sustained permitted churn at a rate the fleet could not absorb.
A lease answers "may I recycle right now?" Nothing answered "has the fleet shed too much capacity this hour?" Multiply your cooldown out before trusting it as a capacity guard: a 120s fleet lease permits 30 recycles/hour fleet-wide. With 6 workers that is replacing the fleet five times an hour, entirely within policy. 13 already generated 20 connection resets, so the permitted ceiling was 2.3x worse than the event that hurt.
Same error as the two earlier lessons, one scope up. Each fix narrowed who may recycle simultaneously (worker -> instance -> fleet); none bounded how often. A cooldown looks like a rate limit because it has a time unit in it, but per-event minimum spacing only caps rate if you multiply it out and compare against absorbable churn.
2. The cooldown gates the recycle even past the hard ceiling
Our recycler has a per-instance cooldown (COOLDOWN_S = 120), shortened to PRESSURE_COOLDOWN_S = 30 when cgroup headroom drops under MIN_HEADROOM_MB = 350, plus a timer-based RSS sample (TICK_S = 45, previously every 50th request). The decision core:
pressure = headroom_mb is not None and headroom_mb < min_headroom_mb
effective_cooldown_s = pressure_cooldown_s if pressure else cooldown_s
if last_recycle_age_s is not None and last_recycle_age_s < effective_cooldown_s:
return False # <-- runs BEFORE the hard-ceiling check
if not pressure and rps > busy_rps and over_limit_age_s < defer_max_s and rss_mb < hard_limit_mb:
return False # hard ceiling short-circuits BUSY DEFERRAL only
return TrueThe hard ceiling short-circuits busy deferral but not the cooldown — deliberately, so an instance can never mass-cull itself. Worst-case over-ceiling dwell is therefore cooldown + tick = 165s.
Our worst event: 710MB against a 650MB hard ceiling (+60MB), on a recycle whose previous same-instance recycle was 150.2s earlier, inside the predicted 120-165s band. It crossed the ceiling, was cooldown-held, and kept allocating until the cooldown cleared and the next 45s tick fired.
The population isolates the hold, non-monotonically
This is the evidence worth copying, because a single event is only consistent with the hypothesis. Non-pressure recycles bucketed by same-instance gap:
| <120s | 0 | — | — |
| 120-140s | 9 | 559.9 | 580 |
| 140-165s | 11 | 597.8 | 710 |
| >165s | 60 | 519.4 | 606 |
RSS peaks in the predicted worst-dwell band (+78MB mean over the >165s bucket) and then falls again for longer gaps, because >165s gaps are routine recycles that were never near the ceiling. The non-monotonicity is what rules out the obvious confound ("longer gap = more accumulation"). Zero recycles below the 120s cooldown, observed min 123.9s: the cooldown itself is airtight.
3. A mechanism we published and then disproved
Worth recording because the wrong version is the intuitive one, and we shipped it before checking.
We first explained the peak-RSS split — pressure recycles topping out at 659MB vs non-pressure at 710MB — as "pressure grants a 4x faster cooldown (30s vs 120s), so those workers are culled sooner and lower." Clean, mechanistic, wrong.
The 30s pressure cooldown was never exercised. Same-instance gaps for all 7 pressure recycles ranged 471.4s to 3945.9s — none within 6x of the 76s bound. It bound nothing, so it contributed nothing to the split. (The tight sub-120s gaps we had also observed were cross-instance, a different primitive; conflating the two is what made the story plausible.)
The real mechanism: pressure fires on instance cgroup headroom, which is independent of any individual worker's RSS, so pressure-culled workers are taken at whatever RSS they happen to hold (mean 593MB, max 659MB). The 710MB is entirely a non-pressure cooldown-hold artifact with no pressure-path contribution. Same conclusion, different cause — and the corrected version narrows the fix: the sole defect is the non-pressure path's missing ceiling-proximity escalation, and the pressure branch needs no change at all.
General form: when two populations differ and one has a faster escalation path, check whether that path was ever on the critical path before crediting it. An escalation that never fired explains nothing, however well its arithmetic lines up.
Related trap on the same telemetry: because the pressure branch fires precisely when headroom < min_headroom_mb, min(headroom) across recycles is a pressure-window statistic by construction and cannot trend. Ours was perfectly bimodal — 7 pressure rows at 204-346MB, 81 non-pressure rows all >=431MB (median 653). We nearly escalated a "headroom collapse" on a 213 -> 204MB day-over-day move that was a 7-sample wobble inside a bounded regime. Track pressure event count instead.
4. The limit jitter floor sets the baseline rate
Across 90 recycles in 23.4h:
- requests-per-worker: min 48, median 817, max 2794
- 47 of 90 (52%) tripped with < 25MB margin over their limit; margin min = 0 (two fired at
rss == limitexactly) - 19 of 90 (21%) recycled after fewer than 400 requests; the worst lived 48 requests
- jittered limits spanned 460-539MB around a base of 500
A worker that recycles after 48 requests never did useful work. Under --preload the master is ~360MB, so a worker drawing a 460MB limit has ~100MB of runway. We had lowered the base 600 -> 500 to stop OOM kills and never re-examined the floor against the preload baseline, so recycle rate became mostly a property of the jitter floor rather than leak velocity. This is also why rate was the wrong alarm: our runbook threshold was ">10/day/instance = leak accelerating" and we had run 4x that every day for a week, so it had trained itself into background noise.
What to do
- Escalate the cooldown on ceiling overshoot. Once
rss_mb >= hard_limit_mb, use the pressure cooldown (or shorter). Two lines, closes the overshoot class, touches only the non-pressure path. Caveat we would have missed: since the pressure cooldown was never exercised, that constant is untested in production — this change routes traffic through it for the first time (~3 events/day for us). Soak it on stage; the original jitter defect only fell out under load test. - Shorten the RSS sample interval adaptively when RSS is within ~10% of the limit. 45s of blind time is the other half of the overshoot.
- Bound recycle rate fleet-wide (token bucket, or refuse to drain when trailing-window recycles exceed N, or when healthy capacity is below a floor).
- Set the jitter floor from the preload baseline, not the base limit. Runway is
limit_floor - boot_rss; assert at startup that it clears a minimum, so lowering the base to fix OOMs cannot silently create churn. - Alarm on worker lifetime, not recycle count.
requests_per_workerp10 catches both defects here and is invariant to fleet size and limit regime, which recycle rate is not. - Correlate recycles against a client-side connection-error canary before hunting a leak. Ours pointed at the recycler in one join. Our leak hunt walked from priority #1 to #5 as evidence arrived: every finding here is response-latency and rate control, not memory volume, and a heap dump would have found nothing actionable.
Meta-lesson, one scope up again: per-worker randomness is not coordination; per-instance coordination is not fleet coordination; fleet coordination is not rate control; and a mitigation path that never fires cannot explain your worst cases. Each time you scale a mutual-exclusion primitive, ask what it permits at full saturation, and which path is slowest when it matters.