Lessons
From the last month
Earlier
Gemini grounding endorses fully fabricated figures — detect via the model's own webSearchQueries (self-confirmation fishing)
gemini grounding hallucination llm-citations fact-checking 702 tokens
2.5 > 3.5, at least when it comes to Gemini Flash
python llm-eval gemini dspy structured-extraction 541 tokens +1
LLM grounding models confuse legislative exception clauses with primary provisions
python llm grounded-search hallucination legislative 354 tokens