arXiv · 2609.34662
Unsupervised Speech Enhancement via Drifting
Abstract
This paper addresses unsupervised speech enhancement in the unpaired setting using drifting methods, where training relies on separate collections of degraded and clean audio without corresponding pairs. While recent drifting approaches enable unpaired training, they do so at a heavy cost: because the objective optimizes only a marginal prior over clean speech, the enhancer gradually loses the input's linguistic content and speaker identity. To fix this, we introduce input-conditioned drifting. We preserve the pull of the clean corpus while re-tethering the output to the degraded input via two mechanisms: an anchor encoder supplies the missing likelihood by pulling toward the input's features, and a key encoder conditions the prior by re-weighting retrieved frames. Neither requires labels or paired data. Using a training-free encoder selection criterion, Word Error Rate on VoiceBank-DEMAND falls to 10.1% (unprocessed: 11.7%), speaker similarity recovers from 0.490 to 0.879, and the recipe transfers in part to dereverberation on WSJ0-REVERB: content improves, rendering quality does not.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Diego Caviedes-Nozal, Liang Xu, Rasmus Kongsgaard Olsson, W. Bastiaan Kleijn. 2026-09-28. Unsupervised Speech Enhancement via Drifting. https://arxiv.org/abs/2609.34662
Cite the original work for its findings. Save a collection to share your selection of sources.