arXiv · 2610.03390
DriftTTS: Few-Step Text-to-Speech Without Distillation via Distribution-Matching Drift
Abstract
Few-step neural text-to-speech models often rely on short- ened diffusion or flow-matching schedules, or on distillation from pretrained multi-step teachers. To avoid these depen- dencies, we present DriftTTS, a few-step mel-spectrogram generator trained without a generative teacher, distillation, or adversarial discrimination. DriftTTS uses a distribution- matching drift objective in a mel-domain feature space defined by raw mels and a frozen masked-autoencoder encoder pretrained on the same LJSpeech training split. On-policy rollout trains the decoder on its own interme- diate states and supports inference up to the trained roll- out depth. On LJSpeech, DriftTTS at NFE=4 achieves 3.87 dB MCD and 3.7% WER, compared with 3.85 dB and 3.4% for Matcha-TTS. In a fully paired blind listen- ing test, DriftTTS obtains 4.18 MOS, compared with 3.96 for Matcha-TTS and 4.22 for ground truth. These results demonstrate competitive few-step synthesis without a pre- trained generative teacher. Code can be found at https: //github.com/BASHLab/driftTTS.git
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mohammad Nur Hossain Khan, Subrata Biswas, Bashima Islam. 2026-10-02. DriftTTS: Few-Step Text-to-Speech Without Distillation via Distribution-Matching Drift. https://arxiv.org/abs/2610.03390
Cite the original work for its findings. Save a collection to share your selection of sources.