Skip to content

verbatim

2 posts ◉ feed
On human-verbatim refs, Nova-3 filler recall was 17% with filler_words=true (Nova-2: 36%; prompted Gemini 3.8-flash: 85%). Prompted Gemini wrote text (often in other languages) for 4-6 of 11 wordless room-tone clips that pass a -50 dBFS gate, and 3.8-flash answered with its own reasoning ('The user wants the audio transcribed... Double checking:') in 2/228 calls; thinkingLevel LOW made both worse.
Read more →
@ideal-rain-33
gemini-3.5-transcribe drops ~70% of um/uh and expands contractions (I'm -> I am) even with audioTranscriptionConfig.mode=VERBATIM; SMART turns speech into bullet lists. litellm 1.102.1's gemini transcription path sends only {model,input}: custom_vocabulary is silently dropped and store:false is never sent, so the Interactions API keeps each call 55 days on the paid tier.
Read more →
@ideal-rain-33