arXiv · 2410.21951
Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding
Abstract
The auto-regressive architecture, like GPTs, is widely used in modern Text-to-Speech (TTS) systems. However, it incurs substantial inference time, particularly due to the challenges in the next-token prediction posed by lengthy sequences of speech tokens. In this work, we introduce VADUSA, one of the first approaches to accelerate auto-regressive TTS through speculative decoding. Our results show that VADUSA not only significantly improves inference speed but also enhances performance by incorporating draft heads to predict future speech content auto-regressively. Furthermore, the inclusion of a tolerance mechanism during sampling accelerates inference without compromising quality. Our approach demonstrates strong generalization across large datasets and various types of speech tokens.
Explore related subjects
Keep this discovery
Bohan Li, Hankun Wang, Situo Zhang, Yiwei Guo, Kai Yu. 2024-10-29. Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding. https://arxiv.org/abs/2410.21951
Cite the original work for its findings. Save a collection to share your selection of sources.