Measured 2026-09-28 on human-verbatim references (AMI meetings + Earnings-21 calls, 3,359 words, 105 fillers):
| model | verbatim WER | filler recall |
|---|---|---|
| gemini-3.8-flash prompted (system prompt: keep fillers, repetitions, false starts) | 5.5% | 85% |
| gemini-3.5-transcribe (Interactions, store:false) | 8.4% | 31% |
| gemini-3.5-transcribe + custom_vocabulary | 8.8% | 27% |
Verbatim mode doesn't exist in practice:
- Interactions
generation_config.transcription_config.mode={type: verbatim}or"verbatim": output identical to default (fillers gone, "I am" for spoken "I'm"). - generateContent
generationConfig.audioTranscriptionConfig.mode: "VERBATIM": identical to default."SMART"rewrote a spoken list into markdown bullets.
litellm 1.102.1 (llms/gemini/audio_transcription/transformation.py): _build_transcription_config reads only language / timestamp_granularities / response_format, so custom_vocabulary= kwargs vanish (verified live: request body was only {model, input}); prompt= raises UnsupportedParamsError. It never sends store:false, and the Interactions API default is store=true (paid tier: 55 days; https://ai.google.dev/gemini-api/docs/interactions-overview). generateContent does not store by default.
Working raw request with vocabulary and no storage:
POST https://generativelanguage.googleapis.com/v1beta/interactions
x-goog-api-key: ...
{"model": "gemini-3.5-transcribe", "store": false,
"input": [{"type": "audio", "mime_type": "audio/ogg", "data": "<b64>"}],
"generation_config": {"transcription_config": {"custom_vocabulary": ["CalVer", "SemVer"]}}}Text is in steps[].content[] where type=="text". usage.total_output_tokens reports 0.
If you need verbatim (fillers, restarts), use a prompted chat model instead, and guard it (see the companion lesson on invented speech and reasoning leaks).