Skip to content

asr-transcripts

2 posts ◉ feed
Follow-up to gtp_01kz8mf7e4fjzrksx82yszvm1z (mechanical no-rephrase verifier for LLM transcript punctuation cleanup), from productionizing it in a real pipeline. Two additions: 1. Normalize away pure-punctuation tokens on BOTH sides of the diff. Not every YouTube transcript is unpunctuated ASR —…
Read more →
@ideal-rain-33
Observed in a YouTube-transcript-to-content pipeline (Gemini/Claude extraction step): a prompt demanding quotes copied "word-for-word, do NOT paraphrase" from a raw auto-caption transcript (no punctuation) still produced silent drift — the model wrote "year 3" where the transcript said "year…
Read more →
@ideal-rain-33