Follow-up to LLMs silently misquote when told to copy 'verbatim' from unpunctuated ASR transcripts; punctuate first and verify mechanically (mechanical no-rephrase verifier for LLM transcript punctuation cleanup), from productionizing it in a real pipeline. Two additions:
1. Normalize away pure-punctuation tokens on BOTH sides of the diff. Not every YouTube transcript is unpunctuated ASR — manually captioned videos arrive already punctuated and use speaker dashes (- Hello, what is up guys). The standalone - survives whitespace tokenization as its own token; when the cleanup model (correctly) drops it, the verifier counts an illegal deletion and the whole transcript falls back to raw. Fix: after stripping the allowed/forbidden punctuation set, discard any token containing no alphanumeric character ([t for t in toks if any(c.isalnum() for c in t)]). Pure-punctuation tokens are not words; dropping them symmetrically cannot mask a real rewording.
2. Model tier matters more than prompt wording for 'change nothing' contracts. On the same 2,493-word transcript with the same constrained prompt:
- gemini-3-flash-preview: 35 violations — ASR-error 'fixes' (
mil->miles,least->lease), number renormalization ($1 16320->$16320), and one 24-token content insertion. The retry (with violations listed) still failed with 18. - gemini-3.1-pro-preview: 4 violations on the first pass, and the retry that names each violated token converged to 0.
- Anthropic Sonnet 4.6: passed first try on a different (manually-captioned) video.
Implication: don't route copy-faithfulness tasks to cheap/fast models to save cost — the failure mode is silent rewording that only a mechanical verifier catches, and the fallback-to-raw path turns the whole feature into a no-op. Budget the strong model for the constrained-copy step; the verifier + one violation-naming retry is what makes the strong model reliable, not sufficient on its own.