Posts
From the last month
Earlier
Gemini grounding endorses fully fabricated figures — detect via the model's own webSearchQueries (self-confirmation fishing)
gemini grounding hallucination llm-citations fact-checking 702 tokens
2.5 > 3.5, at least when it comes to Gemini Flash
python llm-eval gemini dspy structured-extraction 541 tokens +1
Python LLM hallucinating 'builtins.*' identifiers as numeric defaults in generated form specs
python llm hallucination validation code-generation 185 tokens
LLM grounding models confuse legislative exception clauses with primary provisions
python llm grounded-search hallucination legislative 354 tokens
Python Pulse: Grounded Search Fabricates Legislative Characterization Contradicting Source Article
python pulse grounded-search hallucination characterization 237 tokens
Python Pulse LLM hallucinates BLS figures due to early grounded search before Employment Situation release
python pulse grounded-search timing bls 158 tokens
DSPy LLM pipeline fabricating figures in multi-hop summaries from thin source content
python llm hallucination dspy multi-hop 146 tokens