Cloth-rustle / handling noise on a lav-mic talking-head take tempts you into afftdn. Before spending a generation on it, measure three numbers on a representative span, not just SNR:
- SNR = speech RMS - silence RMS.
- LF transient p95 = 95th percentile of summed 20-120 Hz FFT power over ~0.04s hops. This is the impulsive cloth thump, i.e. the thing you actually hear.
- Voiced-frame LSD = RMS log-spectral distance (200-4000 Hz) between original and processed, computed ONLY on frames above a speech threshold. Computing it over the whole span inflates it with the silences the denoiser is supposed to change.
Real measurement on a 14s DJI-Mic-3-into-iPhone take (44.1 kHz AAC ~100 kbps, male speaker, F0 median 113 Hz):
| none | 29.7 | 27.5 | 0.00 |
highpass=f=90 | 31.0 | 26.3 | 0.08 |
highpass=f=100 x2 | 31.2 | 24.0 | 0.17 |
anlmdn | 32.4 | 27.5 | 5.91 |
afftdn=nr=12:nf=-45:tn=1 | 34.1 | 27.8 | 11.26 |
afftdn=nr=20:nf=-40:tn=1 | 37.6 | 27.8 | 13.32 |
The conclusion SNR alone would have hidden: afftdn improves the number you measured and not the artefact you hear. It buys 4-8 dB of noise floor, moves the LF transients by 0.0-0.3 dB (they are impulsive and overlap speech, so a stationary-noise estimator never touches them), and rewrites the voiced spectrum by 11-13 dB. That is the profile of a repair that gets rejected by ear.
A gentle high-pass is the only transparent gain. Pick the corner from the speaker's own F0 distribution (autocorrelation over 60-300 Hz on voiced frames), not from a rule of thumb: at F0 p10 = 95 Hz, a 90 Hz 2-pole high-pass costs 0.08 dB voiced LSD, 0.19 dB of speech level and 0.64 dB of 80-250 Hz voice body while removing 7-10 dB of sub-90 Hz energy. Steeper or higher corners keep taking LF but start eating the fundamental.
One trap while measuring: ffmpeg -v error -af volumedetect prints nothing, because volumedetect reports at info level. Use -hide_banner and leave the log level alone.