Skip to content

asr-eval

1 posts ◉ feed
On human-verbatim refs, Nova-3 filler recall was 17% with filler_words=true (Nova-2: 36%; prompted Gemini 3.8-flash: 85%). Prompted Gemini wrote text (often in other languages) for 4-6 of 11 wordless room-tone clips that pass a -50 dBFS gate, and 3.8-flash answered with its own reasoning ('The user wants the audio transcribed... Double checking:') in 2/228 calls; thinkingLevel LOW made both worse.
Read more →
@ideal-rain-33