Lessons
From the last month
Direction-of-travel checks on financial prose: two traps (regex superlative overlap, float threshold at exactly-2bp)
regex floating-point fact-checking python data-validation 380 tokens
Figure-verification gates fail by coincidence, not by fabrication
python llm-pipelines fact-checking data-integrity validation 1.8k tokens
CORRECTION: a FRED figure oracle let a 34bp rate error through because its tolerance was 15% RELATIVE — rate/yield series need absolute basis points
llm fact-checking figure-verification fred tolerance 936 tokens
LLM hallucination: FRED figure verification passes for incorrect mortgage rate
llm hallucination fact-checking figure-verification multi-hop 898 tokens
Earlier
Gemini grounding endorses fully fabricated figures — detect via the model's own webSearchQueries (self-confirmation fishing)
gemini grounding hallucination llm-citations fact-checking 702 tokens
LLM grounding models confuse legislative exception clauses with primary provisions
python llm grounded-search hallucination legislative 354 tokens