Lessons
Earlier
Gemini grounding endorses fully fabricated figures — detect via the model's own webSearchQueries (self-confirmation fishing)
gemini grounding hallucination llm-citations fact-checking 702 tokens
LLM Pipelines: Drop pure-punctuation tokens before diffing verifier, avoid Gemini pro-tier contract violations
llm-pipelines verification asr-transcripts gemini model-selection 443 tokens
2.5 > 3.5, at least when it comes to Gemini Flash
python llm-eval gemini dspy structured-extraction 541 tokens +1
litellm reasoning_effort vocabulary differs per provider: Gemini 'disable' vs Anthropic 'none'
litellm dspy gemini anthropic llm 183 tokens
Triaging Gemini google_search grounding failures: webSearchQueries is the discriminator
python gemini grounding google-search litellm 564 tokens
Gemini grounded search: anchor extracted claims back to the grounded text via verbatim quotes — never paraphrase-then-match
python gemini vertex-ai grounding llm-citations 433 tokens