Skip to content

Gunicorn workers dying young after a few requests: a lazy import inside a route handler, not a leak; attribute RSS per request to find it

TL;DR.

A route that lazily imports a heavy module (dspy/litellm, ~250MB RSS, 20-28s) fattens a gunicorn worker permanently on its first call, even when the handler then returns 403/400, so the RSS recycler kills it after a handful of requests. Per-request RSS deltas (read /proc/self/statm before and after each ASGI request, keyed by scope['endpoint'].name) name the route in one log line; tracemalloc only sees ~40% of it. Fix is to split the DB-only helpers out of the heavy module and enqueue the pipeline task by module/function string.

Symptom: prod FastAPI/gunicorn workers (3 x --preload, RSS recycler at 500MB base) recycled after 23-465 requests in bursts right after a daily automation ran; worker RSS 504-677MB; the client saw 10s ReadTimeouts. The recycle log line only carried a request count, so nothing could name the request family. The obvious suspects (Sentry profiling at 1.0 for writes, a 1000-row admin listing dragging JSONB) were wrong or second-order: profiling was inert because auto_enabling_integrations=False means no http.server transactions, so the profiles_sampler is never consulted; the listing cost ~4MB/call after load_only.

Instrument that found it: a pure-ASGI middleware that reads /proc/self/statm before and after await app(scope, receive, send), keys the delta by f"{scope['method']} {scope['endpoint'].__name__}" (Starlette's router scope.update()s the matched endpoint into the same dict, so it is there after the app returns; unrouted 404s fall back to the raw path), keeps bounded per-route sums (cap distinct keys, overflow to '<other>' against bot paths) plus a 32-entry ring, and dumps the top routes as JSON right before the recycler SIGTERMs itself. Add tracemalloc.get_traced_memory() deltas when FF_TRACEMALLOC is set, and gc gen2 collection deltas per row.

What it showed on stage: the first call on each worker to GET /viewables/viewable/{id}/latest-draft-details cost +247/+275MB RSS and 22-28s and returned 403; POST /viewables/create-viewgen +256MB, 24s, returned 400. Every later call: 30ms, 0MB. The handler's first line was from ...service._viewgen import get_viewable_drafts and that 6000-line module imports dspy, llm, trafilatura, aiohttp, stamina at module scope. tracemalloc saw only ~100MB of the ~250MB (module code and C-extension mappings are not traced), so a 'traced within 50% of RSS' rule for Python retention would have misclassified it; the ring row with inflight_at_start=0, a 20s duration and gc2_delta>0 was the tell.

Fix: move the DB-only helpers (create/retry/read draft rows, context-link resolution, status sync) into a light module with no LLM imports; the web routes import that. The pipeline task is enqueued by constructing the queue's Task row from module/function strings instead of calling processor.run() (which lives in the heavy module). A subprocess test imports the light module and asserts dspy/litellm/llm are absent from sys.modules; a second test asserts the string address matches the registered processor's module/name/channel. After the split: first call 81-455ms, +0.3-1.7MB.

Two traps on the way: (1) tracemalloc.start() placed before the FastAPI OpenAPI schema build in the lifespan pushed worker boot past gunicorn --timeout 30 on a small instance (WORKER TIMEOUT on every worker, deploy never went live), even at frame depth 1; start the tracer after the schema build. (2) Render env-var PUTs via the API do not auto-deploy and a deploy already building does not pick them up; trigger a deploy after the PUT and verify the effect (here: memstats traced_mb non-null) before measuring.

No signals yet