Search arXiv⌕ Search

arXiv · 2610.05155

UltraM2M: Leveraging Text Transcripts and Mixture Constraints for Weakly-Supervised Speech Enhancement

Abstract

We propose UltraM2M, a weakly-supervised speech enhancement algorithm building upon the recent unsupervised mixture-to-mixture (M2M) algorithm. M2M realizes unsupervised speech enhancement by training deep neural networks on a set of real-recorded noisy-reverberant multi-channel mixture signals to estimate target speech and non-target signals. The two estimated signals are penalized by a so-called mixture-constraint (MC) loss, which constrains them to reconstruct the observed mixture signals. Although shown to be effective, the mixture constraint may be too weak to enable sufficient noise reduction. To deal with this, UltraM2M extends M2M by further leveraging text transcripts of real-recorded mixtures to design an automatic speech recognition (ASR) loss to penalize the estimated speech signal. The ASR loss can be viewed as a form of weak supervision that could help unsupervised enhancement. Evaluation results on the CHiME-4 dataset show the effectiveness of UltraM2M.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Liu, Jiachen, Wu, Fulin, Wang, Zhong-Qiu. 2026-10-04. UltraM2M: Leveraging Text Transcripts and Mixture Constraints for Weakly-Supervised Speech Enhancement. https://arxiv.org/abs/2610.05155

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Personal VAD: Speaker-Conditioned Voice Activity Detection

In this paper, we propose "personal VAD", a system to detect the voice activity of a target speaker at the frame level. This system is useful for gating the inputs to a streaming on-device speech recognition system, such that it only triggers for the target user, which helps reduce the computational cost and battery consumption, especially in scenarios where a keyword detector is unpreferable. We achieve this by training a VAD-alike neural network that is conditioned on the target speaker embedding or the speaker verification score. For each frame, personal VAD outputs the probabilities for three classes: non-speech, target speaker speech, and non-target speaker speech. Under our optimal setup, we are able to train a model with only 130K parameters that outperforms a baseline system where individually trained standard VAD and speaker recognition networks are combined to perform the same task.

eess.AS↗

Textual Echo Cancellation

In this paper, we propose Textual Echo Cancellation (TEC) - a framework for cancelling the text-to-speech (TTS) playback echo from overlapping speech recordings. Such a system can largely improve speech recognition performance and user experience for intelligent devices such as smart speakers, as the user can talk to the device while the device is still playing the TTS signal responding to the previous query. We implement this system by using a novel sequence-to-sequence model with multi-source attention that takes both the microphone mixture signal and source text of the TTS playback as inputs, and predicts the enhanced audio. Experiments show that the textual information of the TTS playback is critical to enhancement performance. Besides, the text sequence is much smaller in size compared with the raw acoustic signal of the TTS playback, and can be immediately transmitted to the device or ASR server even before the playback is synthesized. Therefore, our proposed approach effectively reduces Internet communication and latency compared with alternative approaches such as acoustic echo cancellation (AEC).

eess.AS↗

Multiplexing Neural Audio Watermarks with Adaptive Routing

Audio watermarking supports speech authenticity verification. We study whether combining heterogeneous watermarks can retain at least one detectable provenance signal when their failure modes differ. This Any-survival objective concerns complementary evidence retention, not simultaneous survival of all constituent marks or recovery of all payloads. To our knowledge, this is the first systematic study of neural audio watermark multiplexing, with a scoped benchmark covering five released systems, parallel and sequential baselines, and 14 evaluation conditions. The benchmark shows complementary failure modes, but naive composition does not reliably turn them into system-level robustness under the Any-survival objective. We therefore formulate multiplexing as watermark allocation and study perceptual-adaptive time-frequency multiplexing (PA-TFM), a training-free routing method, and MaskNet, a learned time-domain router for separately trained systems with native detectors. We use the five-system scoped benchmark to characterize multiplexing behavior and the AudioSeal-PerTh pair as a representative heterogeneous pair for adaptive watermark routing. Compared with direct parallel composition, MaskNet improves average TPR@1%FPR from 0.75 to 0.88 and SNR from 15.20 dB to 25.36 dB; compared with PA-TFM, it gives a clear system-level robustness gain under the Any-survival objective while retaining similar high-fidelity behavior. A SpeechTokenizer-aware case study further improves average TPR@1%FPR to 0.91 and SpeechTokenizer robustness from 0.20 to 0.60, showing that channel-adapted constituents can enter the same routing framework.

eess.AS↗