Context: product needs a verbatim transcript (um/uh, repeats, false starts kept) of 30 s to 5 min phone recordings, fast enough for a synchronous UI. Scanned the 2026-09 STT market after measuring Gemini flash (prompted; ~85% filler recall, 2.5 s p50) and Deepgram Nova-2/3 (0.2 s p50; Nova-3 kept 17% of fillers even with filler_words=true).
Where to look. Artificial Analysis AA-WER and the HF Open ASR leaderboard run Whisper-style normalization that strips fillers before scoring, so Deepgram Nova-3 looks fine there (5.2% AA-WER) while dropping most disfluencies. The one public benchmark that scores fillers, repetitions, cut-offs and vocal sounds as typed metrics is Nyra Labs' verbatim speech benchmark (MIT evaluator + cached predictions for 15 systems): https://nyra-labs.com/research/nyra-verbatim-speech-benchmark and https://github.com/nyrahealth/nyra_verbatim_speech_benchmark. English filler F1 / vWER as of 2026-06-24: CrisperWhisper 2.0 Pro 95.7 / 3.0; ElevenLabs Scribe v2 95.5 / 3.2; MAI-Transcribe-1.5 95.8 / 3.6; Inworld STT 95.2 / 4.0; AssemblyAI Universal-3 Pro 90.8 / 4.5; xAI Grok 84.4 / 4.7; Deepgram Nova-3 45.7 / 6.3; NVIDIA Canary 1B v2 26.7 / 9.0; Whisper large-v3 9.4 / 10.2; Cartesia Ink-Whisper 6.0 / 10.4. Gemini chat models aren't scored anywhere; measure them yourself.
ElevenLabs Scribe v2 is the hosted option with both verbatim numbers and speed (Artificial Analysis speed factor 86x, i.e. ~0.7 s per audio minute; $0.22/hr). Verbatim is the default: no_verbatim=true is the opt-out ("will not have any filler words, false starts and non-speech sounds"), per https://elevenlabs.io/docs/api-reference/speech-to-text/convert. Send tag_audio_events=false unless you want "(laughter)" in the text; keyterms (list, +20% price) for names. Gotchas: training on your audio is ON by default for non-enterprise accounts (toggle in profile > Terms and privacy > Data use); zero retention (enable_logging=false) is enterprise-only; litellm's ElevenLabs transcription route only knows scribe_v1 and passes none of these params, so call it with a plain multipart POST (xi-api-key header, fields model_id=scribe_v2, file).
Cloudflare Workers AI hosts exactly four ASR models (checked 2026-09-29): @cf/openai/whisper, @cf/openai/whisper-large-v3-turbo (~1 MB practical request limit; CF's own tutorial chunks at 1 MB), @cf/deepgram/nova-3 ($0.0052/min HTTP vs $0.0043 direct; mip_opt_out still defaults off), and @cf/deepgram/flux (turn detection). If your backend isn't on Cloudflare it is just another HTTPS hop with a worse model or a 20% markup. AI Gateway only proxies third-party STT with logging/fallback. https://developers.cloudflare.com/workers-ai/platform/pricing/
Whisper family (Groq, Fireworks, CF) is cheapest but off the table for verbatim: ~9% filler F1 and a ~40% false-positive rate on silent audio (https://arxiv.org/pdf/2609.04561). Parakeet-TDT is claimed verbatim in one GitHub issue only; its sibling Canary scores 26.7% filler F1, so treat that as unverified.