Context: FastAPI + litellm app, tests replay VCR cassettes (vcrpy) for Gemini calls. Cassette matching was match_on=['method','uri'] (the default-ish setup that ignores the body), and every chat completion (writeup, topic extraction, vibe pick, STT prompt) goes to the same URI .../v1beta/models/gemini-x:generateContent.
Symptom: after adding suggest_thanks() (one more complete_json call inside create_yap, the shared constructor for request yaps), about 20 existing tests across unrelated modules failed with schema/key errors: each cassette's responses are consumed in order per URI, so the new call ate the response recorded for the next call (topic extraction got the vibe JSON, etc.). Nothing in the failing tests mentioned the new feature.
Why re-recording is wrong: it is dozens of cassettes, some needing real keys, and the next feature does it again.
Fix that held: a real config switch ([thanks].suggest_enabled, default true, false in the devtest config) checked at the top of the new function (returns None, the same fail-closed path the feature already has for model failures). One new test flips the flag on via a fixture and records its own cassette. All old cassettes replay unchanged.
Alternatives considered: matching on body (fragile: prompt edits break every cassette; also litellm's body includes tool schemas) or allow_playback_repeats (hides ordering bugs). The flag is also a product knob (cost control), which is why it was acceptable.
General rule: with URI-only matching and a single-endpoint model gateway, any new model call in a shared path needs either its own URI, or a test-config kill switch, before you touch existing cassettes.