arXiv · 2610.08276
Voice Anonymization Made Simple: Training-Free Anonymization with Projected Classifier-Free Guidance
Abstract
Voice anonymization aims to conceal speaker identity while preserving linguistic and paralinguistic information, yet many existing methods require dedicated training. We propose a training-free, two-stage anonymization method for pretrained flow-matching voice conversion systems. In the condition-construction stage, the source content representation is regenerated under a randomly selected speaker context to reduce residual source-speaker information while preserving linguistic content. In the speech-generation stage, the source speech is used to identify the generation direction associated with the original speaker, and the source-aligned component is suppressed while useful content guidance is retained. The proposed method modifies only inference-time conditioning and guidance, without updating any pretrained parameters. Experiments on VoicePrivacy 2026 Track 1 show that the two-stage method improves speaker privacy while maintaining competitive word error rate and emotion recognition performance.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xiang Shi, Han Zhu, Ming Li, Xiaoxiao Miao. 2026-10-06. Voice Anonymization Made Simple: Training-Free Anonymization with Projected Classifier-Free Guidance. https://arxiv.org/abs/2610.08276
Cite the original work for its findings. Save a collection to share your selection of sources.