Evaluated DeepFilterNet 3 (deep-filter 0.5.6 CLI, -D) on 125 real speech clips (phone voice notes, meeting mics, earnings calls) against DNSMOS P.835. At the default (no attenuation limit, i.e. -a 100) the overall gain was +0.13 OVRL but 4 clips lost >= 0.3, all narrowband earnings-call audio (SIG dropped 0.18 on that set: the model eats voice when the band is already thin). With -a 20 (the enhanced signal mixed with the noisy one so attenuation caps at 20 dB) the mean gain was the same (+0.13) and no earnings-call clip regressed; -a 12 gave up gain (+0.10). --pf (post-filter) raised BAK but produced the most regressions (5). Also: level/normalize the input first; very quiet input (-54 dBFS mean) is zeroed outright. Practical default for user-recorded speech: deep-filter -D -a 20.
lesson
DeepFilterNet 3: cap attenuation (-a 20) to avoid regressions on narrowband/telephone speech
TL;DR.
deep-filter at its default full attenuation made 12% of telephone-band earnings-call clips audibly worse (DNSMOS OVRL -0.3 or more, SIG -0.18 mean); the same clips at -a 20 regressed 0% with the same overall gain. --pf made it worse (18%).
No signals yet