arXiv · 2310.03538
Latent Filling: Latent Space Data Augmentation for Zero-shot Speech Synthesis
Abstract
Previous works in zero-shot text-to-speech (ZS-TTS) have attempted to enhance its systems by enlarging the training data through crowd-sourcing or augmenting existing speech data. However, the use of low-quality data has led to a decline in the overall system performance. To avoid such degradation, instead of directly augmenting the input data, we propose a latent filling (LF) method that adopts simple but effective latent space data augmentation in the speaker embedding space of the ZS-TTS system. By incorporating a consistency loss, LF can be seamlessly integrated into existing ZS-TTS systems without the need for additional training stages. Experimental results show that LF significantly improves speaker similarity while preserving speech quality.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jae-Sung Bae, Joun Yeop Lee, Ji-Hyun Lee, Seongkyu Mun, Taehwa Kang, Hoon-Young Cho, Chanwoo Kim. 2023-10-05. Latent Filling: Latent Space Data Augmentation for Zero-shot Speech Synthesis. https://arxiv.org/abs/2310.03538
Cite the original work for its findings. Save a collection to share your selection of sources.