DiffVQE2: An Efficient Low-delay Diffusion Model for Acoustic Echo and Noise Control
Hands-free communication devices and speakerphones are inherently affected by acoustic echo and background noise. To mitigate these impairments, end-to-end discriminatively trained neural networks have emerged as the best-performing approach in research and deployment. While recent advancements in generative methods have provided remarkable results for various speech enhancement tasks, diffusion-based acoustic echo control (AEC) research is still restricted to non-causal, utterance-level processing, thereby not widely applicable in practice. With this work, we are the first to propose low-delay (i.e., causal) diffusion-based joint AEC and noise control models DiffVQE2 / DiffVQE2-S, excelling the so-far state of the art DeepVQE / DeepVQE-S models in multiple objective metrics, and, most importantly, in subjective MOS, respectively. In addition, our models are less complex. Furthermore, we show that a limited lookahead applied to the efficient DiffVQE2-S model allows for an even higher performance. These results have been obtained on the ICASSP 2023 AEC Challenge blind test set. We claim the first streaming-capable, diffusion-based acoustic echo and noise control that excels state-of-the-art discriminative approaches.