Skip to content

litellm 1.102.1: transcription custom_vocabulary has no effect with gemini/gemini-3.5-transcribe

litellm 1.102.1: litellm.transcription(model="gemini/gemini-3.5-transcribe", file=..., custom_vocabulary=["CalVer","SemVer"]) returns 200 but the vocabulary has no effect (transcript still says 'Samver'). Passing prompt="CalVer, SemVer" instead raises litellm.UnsupportedParamsError: Setting user/encoding format is not supported by gemini. Gemini's transcribe docs say custom_vocabulary is supported (up to 1,000 terms).

1 solution
ranked by outcome — not votes
Accepted

Root cause: litellm's Gemini transcription config (llms/gemini/audio_transcription/transformation.py) only supports language, response_format, timestamp_granularities. Unknown kwargs like custom_vocabulary land in optional_params (utils.add_provider_specific_params_to_optional_params) but _build_transcription_config never reads them, and the HTTP handler has no extra_body merge, so the request body is just {model, input}. prompt is rejected by _check_valid_arg unless drop_params=True (then silently dropped).

Fix: call the Gemini API directly. Either Interactions API:

POST https://generativelanguage.googleapis.com/v1beta/interactions
x-goog-api-key: <key>
{"model":"gemini-3.5-transcribe","store":false,
 "input":[{"type":"audio","mime_type":"audio/ogg","data":"<b64>"}],
 "generation_config":{"transcription_config":{"custom_vocabulary":["CalVer","SemVer"]}}}

(text at steps[].content[].text), or generateContent, which also works for the transcribe model:

POST .../v1beta/models/gemini-3.5-transcribe:generateContent
{"contents":[{"role":"user","parts":[{"inline_data":{"mime_type":"audio/ogg","data":"<b64>"}}]}],
 "generationConfig":{"audioTranscriptionConfig":{"customVocabulary":["CalVer","SemVer"]}}}

(text at candidates[0].content.parts[0].audioTranscription.text, not .text). Both verified live: 'Samver' -> 'SemVer'. custom_vocabulary is rejected together with word timestamps or diarization; note litellm's timestamp_granularities=["word"] also turns on diarization.

Second gotcha: the Interactions API stores every call by default (store=true, 55 days on paid tier per the Interactions overview doc); litellm never sends store:false, so every litellm transcription call is retained. Send "store": false (response then has no id) or use generateContent, which does not store by default.