Gemini with LiteLLM/DSPy: "source_name" attribution mismatch from grounded search metadata
News pipeline on Gemini grounded search via litellm 1.93.0 (Phase A: litellm.completion with the googleSearch tool -> grounded text + groundingChunks/groundingSupports from _hidden_params['vertex_ai_grounding_metadata']; Phase B: a dspy.Predict extraction, dspy 3.2.1, emitting headline/why_it_matters/source_name/source_quote). URLs come from the grounding chunks, never the LLM.
Auditing published output, source_name frequently disagrees with the linked domain: "U.S. Census Bureau" -> a private bank's markets blog, "CME FedWatch Tool" -> cbsnews.com, "U.S. Treasury" -> bankrate.com, "U.S. Strategic Petroleum Reserve" -> a public radio station, "Energy market reports" -> latimes.com. Links all resolve 200 and every cited figure is present in the linked article, so link-liveness and figure verification pass clean. No error is raised.
The obvious read is 'the model is naming the wrong thing, derive the outlet from the resolved host instead.' Built that and it was wrong — it flagged correct output and would have broken the product's intended attribution style. Then tried a known-outlet allowlist to catch only the invented names; it does not fire, because the bad values are real famous institutions that pass any 'is this a real organization?' test. Neither name-vs-domain comparison nor entity validation separates the good cases from the bad ones, and the bad ones are a minority hidden inside a majority of correct output.
Name-vs-domain is the wrong predicate. Origin-vs-figure is the right one.
First, the disagreement is usually correct and often deliberate. Our Phase B signature reads:
source_name (the primary outlet or releasing agency)That or releasing agency is intentional. .gov endpoints 403 datacenter IPs, so the grounding chunk for a government release is nearly always a news outlet reporting on it. [U.S. Bureau of Labor Statistics](theguardian.com/...) is the desired output, not a defect — the reader learns who produced the number and gets a reachable link. Same for [CME FedWatch Tool](cbsnews.com/...): CBS's own text reads "According to the CME Group's FedWatch tool, there's around a 35% chance." Deriving the outlet from the host destroys real information.
The defect is narrower: the named body must actually produce the cited figure. Phase B reads only the grounded text, so it infers the origin from prose and gets it wrong whenever the story's headline sounds institutional but its numbers come from a private survey. Two classes:
- Wrong body — a real institution that does not publish that series.
"U.S. Treasury"credited with 30yr 6.75% / 15yr 6.10%, which are a rate-comparison site's own survey averages."U.S. Strategic Petroleum Reserve"credited with $4.11/gal (an auto club's survey) and >$100/bbl (an exchange settle) — the SPR is a stockpile and publishes neither. The reliable trap: rate and price SURVEYS (Bankrate, Freddie Mac PMMS, AAA, NAR) and exchange settles getting attributed to a government agency because the headline mentions one. - Not an entity — a noun phrase nobody can attribute:
"Energy market reports","U.S. Energy Markets","U.S. gas stations".
The check. For each story, pair the figure in why_it_matters with source_name and ask whether that body publishes that series. Class 2 is catchable statically (is the string an organization?). Class 1 needs the pairing — an allowlist of institution names cannot see it, because the name is legitimate and only the pairing is false. Cheapest implementation: have Phase B emit the figure's producer as a separate field with the instruction that it must be the organization that published the number, then reject when it disagrees with source_name or is not an attributable body.
Grep your fixtures before you start. Our recorded VCR cassette for the extraction call had been green for weeks containing "source_name": "The Federal Reserve" and "U.S. gas stations", asserted as correct by a passing test. If you have fixtures for this call, the defect is probably already frozen into your test data.
Unrelated but co-located: grounded summaries of nonprofit/public-media pages absorb the site's fundraising banner into the story body — an oil-price story whose 'what this means for you' text asked the reader to start a monthly donation. Strip donation/subscribe boilerplate from grounded text before extraction.