Context: product needs a strictly verbatim transcript (um/uh, repeats, false starts). Eval 2026-09-28, 78 clips, identical ogg/opus bytes to every model; scorer keeps fillers (public leaderboards like Artificial Analysis and HF Open ASR strip fillers before WER, so they can't rank this).
Deepgram: model=nova-3&filler_words=true&smart_format=false&punctuate=true turned "Uh I'm Jeanne-Oui. Um uh my role is..." into "Yeah. I'm Janvi. My role is..."; same with/without language, mip_opt_out, punctuate. nova-2 with the same params: "Yeah. Uh, I'm Janvi. I'm, um, my role is... are, are, uh,". Deepgram's models page itself says nova-2 is recommended for "filler word identification". Nova-2 takes keywords= not keyterm=; keyword boosting of an entity list made "and" -> "IND" 7 times, so skip it. Don't use numerals=true just to align number formats with a digit-writing baseline: Deepgram then mixes forms ("2nd quarter", "20 20"), which Whisper's number normalizer mangles.
Gemini chat STT (3.8-flash / 3.7-flash, system prompt demanding verbatim):
- Invented speech: 11 phone takes of room tone (peak above -50 dBFS, so a dB gate passes them). Every run, Gemini produced text for 4-6: "Настя, играем! Оп! Оп!", "Hayırlı işler. Kolay gelsin.", "Haben Sie ein Telefonbuch?", "So, wait, is that a cat or a dog?". Deepgram Nova-2 produced nothing on all 11. A dB silence gate is not enough; use a VAD or a second-model speech check, and flag rather than drop to avoid losing real quiet speech.
- Reasoning leak: 2/228 calls of 3.8-flash (default thinking) returned "The user wants the audio transcribed word for word.\nAudio:\n- "..."\n\nDouble checking: ..." as the answer text (not a thought part). thinkingLevel LOW: 5/76. Adding a key-terms list to the user turn: 1/76 ("Key terms provided: ..."). 3.7-flash: 0/228. Guard with a regex on meta phrases (the user wants|double[- ]check|key terms provided|the audio contains) and treat a hit as a failed attempt.
- Latency p50 ~2.5 s, p95 6-9 s vs Deepgram p50 0.17 s.
Also: all Gemini models rewrote spoken "gonna" as "going to".