LiteLLM: gemini/gemini-3.5-transcribe "response.choices[0].message.content" empty, transcript in raw response
Calling litellm.completion(model="gemini/gemini-3.5-transcribe", messages=[{role: user, content: [{type: file, file: {file_data: data_url}}]}]) succeeds (HTTP 200, prompt tokens billed, response_cost set) but choices[0].message.content is empty and completion_tokens == 0. The raw generateContent response puts the transcript in candidates[0].content.parts[0].audioTranscription.text, a part type litellm's chat transformation does not read, so it looks like a silent recording rather than an integration error.
Observed on litellm 1.102.1, Gemini Developer API, September 2026. Same clip through the chat models (gemini-3.8-flash / 3.7-flash / 3.5-flash with a verbatim system prompt) transcribed fine, which made the empty result look like a model quirk instead of a parsing gap.
Route by model: litellm.transcription(model="gemini/gemini-3.5-transcribe", file=("speech.ogg", audio_bytes, "audio/ogg"), api_key=...) goes through litellm's Gemini Interactions-API transcription transform and returns response.text plus usage.input_token_details.audio_tokens and _hidden_params["response_cost"] (a 6 s clip: 1.7 s, $0.00049). VERBATIM mode (keeps um/uh, repeats, false starts) is the API default, so no prompt or config is needed.
Do not send chat models down litellm.transcription() as a fallback: gemini/gemini-3.8-flash via that route returned a "Cleaned version:" with fillers removed, because there is no system prompt on that path. Keep chat-model fallbacks on litellm.completion() with an explicit verbatim system prompt, and pick the route by "transcribe" in model.
Also: litellm.completion() warns No text in user content. Adding a blank text for a file-only message; harmless for chat models, but another sign you are on the wrong path for the transcribe model.