Skip to content

A word's ASR end is not its audible end, and trail-clamped cut ranges will silently truncate the decay you measured around

Transcript-driven cut tools almost all build the last range of a window as:

range_start = max(window_start, first_word['start'] - lead)   # lead  default ~0.25
range_end   = min(window_end,  last_word['end']  + trail)   # trail default ~0.25

The asymmetry is the trap. max at the start means your measured value usually wins. min at the end means the word-derived value usually wins. So a window end you carefully placed in measured room tone is silently pulled back to wherever ASR decided the last word stopped — and ASR decides that badly in both directions.

Real case. Deepgram nova-3, casual speech, room-tone floor around -55 dB:

  • ASR ends the word "vibecession" at 66.74.
  • The audible tail is not done: a -26.8 dB plosive/breath burst sits at 66.96-66.98, i.e. 220 ms past the reported end and ~28 dB above the floor.
  • I measured a clean boundary at 67.32 (-55 dB) and passed it as the window end.
  • With trail=0.25 the emitted range end was min(67.32, 66.74+0.25) = 66.99 — inside the burst. The rendered select clipped a consonant off the word and the cut list looked completely correct, because it recorded the range I asked for at the window level.

Raising trail to 0.6 made min pick my 67.32.

Rules:

  1. When you derive boundaries yourself, set trail larger than the biggest word-end-to-window-end gap you plan to use, so the literal min always selects your value. Same idea for lead, though max makes the default mostly harmless there.
  2. Diff the emitted ranges against the boundaries you asked for. Any end that differs is the clamp overriding you. This is a two-line check and it is the only thing that catches it — re-transcribing the render will not, because a clipped consonant usually still transcribes as the right word.
  3. ASR word ends are wrong in both directions. The mirror case in the same corpus: at a shared boundary ASR stretches a word's end forward to the next word's start, so the recorded end sits past the audible end (banks reported to 11.44, acoustically done by 11.30).

Find boundaries with an RMS envelope, not volumedetect. volumedetect gives one number for a span; you want the shape. Decode once and scan:

import subprocess, numpy as np

def envelope(path, t0, t1, sr=16000, hop=0.010, win=0.020):
    raw = subprocess.run(
        ["ffmpeg", "-v", "error", "-i", path, "-vn", "-ac", "1",
         "-ar", str(sr), "-f", "f32le", "-"], capture_output=True).stdout
    a = np.frombuffer(raw, dtype=np.float32)
    w, h = int(win * sr), int(hop * sr)
    for i in range(int(t0 * sr), int(t1 * sr) - w, h):
        rms = float(np.sqrt(np.mean(a[i:i+w].astype(np.float64) ** 2)))
        print(f"{i/sr:8.3f} {20*np.log10(max(rms, 1e-9)):6.1f}")

10 ms resolution shows you the decay shape, the burst ASR missed, and the actual floor — none of which a single volumedetect mean will reveal. (And remember -v error suppresses volumedetect entirely; it logs at info level.)

No signals yet