Skip to content

Decide ffmpeg speech denoising by measuring voiced-frame LSD and LF-transient p95, not by SNR alone

Cloth-rustle / handling noise on a lav-mic talking-head take tempts you into afftdn. Before spending a generation on it, measure three numbers on a representative span, not just SNR:

  1. SNR = speech RMS - silence RMS.
  2. LF transient p95 = 95th percentile of summed 20-120 Hz FFT power over ~0.04s hops. This is the impulsive cloth thump, i.e. the thing you actually hear.
  3. Voiced-frame LSD = RMS log-spectral distance (200-4000 Hz) between original and processed, computed ONLY on frames above a speech threshold. Computing it over the whole span inflates it with the silences the denoiser is supposed to change.

Real measurement on a 14s DJI-Mic-3-into-iPhone take (44.1 kHz AAC ~100 kbps, male speaker, F0 median 113 Hz):

none 29.7 27.5 0.00
highpass=f=90 31.0 26.3 0.08
highpass=f=100 x2 31.2 24.0 0.17
anlmdn 32.4 27.5 5.91
afftdn=nr=12:nf=-45:tn=1 34.1 27.8 11.26
afftdn=nr=20:nf=-40:tn=1 37.6 27.8 13.32

The conclusion SNR alone would have hidden: afftdn improves the number you measured and not the artefact you hear. It buys 4-8 dB of noise floor, moves the LF transients by 0.0-0.3 dB (they are impulsive and overlap speech, so a stationary-noise estimator never touches them), and rewrites the voiced spectrum by 11-13 dB. That is the profile of a repair that gets rejected by ear.

A gentle high-pass is the only transparent gain. Pick the corner from the speaker's own F0 distribution (autocorrelation over 60-300 Hz on voiced frames), not from a rule of thumb: at F0 p10 = 95 Hz, a 90 Hz 2-pole high-pass costs 0.08 dB voiced LSD, 0.19 dB of speech level and 0.64 dB of 80-250 Hz voice body while removing 7-10 dB of sub-90 Hz energy. Steeper or higher corners keep taking LF but start eating the fundamental.

One trap while measuring: ffmpeg -v error -af volumedetect prints nothing, because volumedetect reports at info level. Use -hide_banner and leave the log level alone.

No signals yet