Skip to content

vcrpy 8.3.0 httpx 0.28 LLM calls via litellm: CannotOverwriteExistingCassetteException in record_mode 'none' with new unrecorded calls

vcrpy 8.3.0 with httpx 0.28 (LLM calls via litellm to one Gemini generateContent endpoint), cassettes in record_mode 'none'. I added one extra best-effort model call early in a code path (errors are caught and it returns an empty result). The tests covering that new call stayed green. But six unrelated older tests, whose cassettes I hadn't touched, started failing further down their flow with: CannotOverwriteExistingCassetteException: Can't overwrite existing cassette (...) in your current record mode (<RecordMode.NONE: 'none'>). No match for the request (<Request (POST) https://generativelanguage.googleapis.com/v1alpha/models/...:generateContent>) was found. Found 2 recorded request(s) matching (...) but they have already been consumed. I expected the new, unrecorded request to fail on its own (and be swallowed by my try/except), leaving the recorded responses for the calls that were recorded.

1 solution
ranked by outcome — not votes
Accepted

vcrpy's default match_on is ['method', 'scheme', 'host', 'port', 'path', 'query']. The request body is not part of it. Every call to one LLM endpoint looks identical to the matcher, so recorded interactions are replayed strictly in order. A newly inserted call does not miss. It matches and consumes the next recorded response (one meant for a later, different prompt), and each later call shifts down by one. The last call then finds its matches already been consumed and raises. Because the new call is best-effort, it quietly gets a wrong-but-parseable response, and the failure surfaces far away in unrelated assertions.

Fix 1: re-record what shifted

Run the whole suite once to find the victims (tests where the new call now sits before recorded ones), then rewrite just those cassettes and read the new outputs:

pytest -p no:cacheprovider --record-mode=rewrite \
  'tests/test_a.py::test_one' 'tests/test_b.py::test_two'

Fix 2: make LLM matching body-aware

So a new call misses (and your error handling runs) instead of stealing a response:

import json

def llm_prompt(r1, r2):
    # Compare a stable part of the body; keys/timestamps must not be in it.
    def key(r):
        body = json.loads(r.body or b'{}')
        return json.dumps(body.get('systemInstruction'), sort_keys=True)
    assert key(r1) == key(r2)

@pytest.fixture(scope='module')
def vcr(vcr):
    vcr.register_matcher('llm_prompt', llm_prompt)
    return vcr

mark_vcr = pytest.mark.vcr(match_on=['method', 'uri', 'llm_prompt'])

Be aware that swallowing LLM errors hides this: the stolen response parses fine for the new call, so the symptom shows up only in whichever later call runs out.