arXiv · 2609.37007
RVQ Position Aware Speculative Decoding for On Device Text to Speech
Abstract
Autoregressive decoding (AR) with Transformer models is memory bandwidth bound at single stream inference, the typical deployment regime for on device text to speech (TTS). Real time streaming with Qwen3-TTS requires more than 200 sequential model calls per second, dominated by the inner loop MultiCodeDecoder that emits the 15 residual vector quantization (RVQ) codes per 80 ms audio frame. We propose RVQ position aware speculative decoding for the MultiCodeDecoder, attaining 2.47 accepted tokens per model call at $5\times10^{-4}$ percent added parameters and 10 to 20 percent per round speculation/verification overhead, reducing real time synthesis from 200 to 88 sequential model calls per second. The scheme is distributionally lossless under the deployed top-k sampling, and WER parity with the original system is consistent with this guarantee. We deliver 2 to 2.2x speedup for RVQ token generation with Qwen3-TTS 0.6B on recent iPhone and Apple Silicon Mac devices.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Berkin Durmus, Eduardo Pacheco, Zach Nagengast, Atila Orhon. 2026-09-29. RVQ Position Aware Speculative Decoding for On Device Text to Speech. https://arxiv.org/abs/2609.37007
Cite the original work for its findings. Save a collection to share your selection of sources.