Ruiqiang Huang trained a Background Voice Cancellation model on top of DPDFNet-8 — 20 epochs on room-acoustic call center data (Libri2Mix + pyroomacoustics), raising SI-SDR improvement from +5.7 dB to +8.3 dB and reducing background voice leakage by 5.5 dB over the pretrained baseline. Evaluated at both native 16 kHz wideband VoIP and 8 kHz G.711 narrowband telephony.
| Metric | Pretrained 16k | Fine-tuned 16k | Pretrained 8k | Fine-tuned 8k |
|---|---|---|---|---|
| Near SI-SDR ↑ | +10.68 dB | +13.31 dB | +7.86 dB | +11.15 dB |
| Far SI-SDR ↓ | −20.65 dB | −26.17 dB | −21.96 dB | −26.24 dB |
| SI-SDR improvement | +5.68 dB | +8.31 dB | +2.87 dB | +6.15 dB |
| PESQ-WB ↑ | 1.749 | 2.244 | 1.527 | 2.036 |
| STOI ↑ | 0.876 | 0.916 | 0.812 | 0.910 |
| Scenario | No processing | Pretrained | Fine-tuned ep20 |
|---|---|---|---|
| Drama 9 dB (~5 m) | 1.68 / 1.21 / 1.30 | 3.12 / 3.88 / 2.81 | 3.37 / 4.04 / 3.10 |
| Call center 1.0 m | 3.40 / 3.09 / 2.61 | 3.43 / 4.08 / 3.15 | 3.53 / 4.12 / 3.24 |
| Call center 0.5 m | 3.41 / 2.94 / 2.54 | 3.38 / 3.84 / 3.00 | 3.45 / 3.99 / 3.11 |
| Scenario | Pretrained | Fine-tuned ep20 | SIG drop (pretrained) | SIG drop (ep20) |
|---|---|---|---|---|
| Drama 9 dB | 3.16 / 3.89 / 2.85 | 3.24 / 4.01 / 2.96 | +0.04 | −0.12 |
| Call center 1.0 m | 2.98 / 4.05 / 2.74 | 3.45 / 4.10 / 3.17 | −0.45 | −0.07 |
| Call center 0.5 m | 2.99 / 3.76 / 2.60 | 3.33 / 3.96 / 3.00 | −0.38 | −0.12 |
Author: Ruiqiang Huang ·
Audio generated by scripts/generate_web_samples.py --n 5 --seed 42 ·
RNNoise via librnnoise (xiph.org) ·
DPDFNet-8 (3.54M params) · Config: configs/dpdfnet8_general_voip.yaml
Real recordings (not simulated) — same near-end talker captured against four different backgrounds, 20 s each. Both bands come from the source recordings; the 8 kHz set is band-limited to 4 kHz before processing. No clean reference exists for these, so there is no target track. Follows the band toggle above.