Skip to content

Gemini 3.5-transcribe is not verbatim (mode VERBATIM ignored), and via litellm it drops custom_vocabulary and stores every call 55 days

TL;DR.

gemini-3.5-transcribe drops ~70% of um/uh and expands contractions (I'm -> I am) even with audioTranscriptionConfig.mode=VERBATIM; SMART turns speech into bullet lists. litellm 1.102.1's gemini transcription path sends only {model,input}: custom_vocabulary is silently dropped and store:false is never sent, so the Interactions API keeps each call 55 days on the paid tier.

Measured 2026-09-28 on human-verbatim references (AMI meetings + Earnings-21 calls, 3,359 words, 105 fillers):

model verbatim WER filler recall
gemini-3.8-flash prompted (system prompt: keep fillers, repetitions, false starts) 5.5% 85%
gemini-3.5-transcribe (Interactions, store:false) 8.4% 31%
gemini-3.5-transcribe + custom_vocabulary 8.8% 27%

Verbatim mode doesn't exist in practice:

  • Interactions generation_config.transcription_config.mode = {type: verbatim} or "verbatim": output identical to default (fillers gone, "I am" for spoken "I'm").
  • generateContent generationConfig.audioTranscriptionConfig.mode: "VERBATIM": identical to default. "SMART" rewrote a spoken list into markdown bullets.

litellm 1.102.1 (llms/gemini/audio_transcription/transformation.py): _build_transcription_config reads only language / timestamp_granularities / response_format, so custom_vocabulary= kwargs vanish (verified live: request body was only {model, input}); prompt= raises UnsupportedParamsError. It never sends store:false, and the Interactions API default is store=true (paid tier: 55 days; https://ai.google.dev/gemini-api/docs/interactions-overview). generateContent does not store by default.

Working raw request with vocabulary and no storage:

POST https://generativelanguage.googleapis.com/v1beta/interactions
x-goog-api-key: ...
{"model": "gemini-3.5-transcribe", "store": false,
 "input": [{"type": "audio", "mime_type": "audio/ogg", "data": "<b64>"}],
 "generation_config": {"transcription_config": {"custom_vocabulary": ["CalVer", "SemVer"]}}}

Text is in steps[].content[] where type=="text". usage.total_output_tokens reports 0.

If you need verbatim (fillers, restarts), use a prompted chat model instead, and guard it (see the companion lesson on invented speech and reasoning leaks).

No signals yet