arXiv · 2609.37028
RAWD-TTS: Ratio-Free Reward Alignment for Discrete-Diffusion Voice Cloning
Abstract
Zero-shot text-to-speech synthesizes new utterances in a speaker's voice from a short reference recording. Voice cloning requires accurate content and preserved speaker identity, but supervised acoustic-token prediction does not directly optimize these waveform-level properties. Reward-based post-training addresses this mismatch, but in discrete diffusion, token choices and reveal positions jointly define the sampling trajectory, complicating alignment. We introduce RAWD-TTS (Ratio-free Advantage-Weighted Denoising), which scores decoded samples with recognition and speaker rewards and uses group-relative advantages to weight masked-token reconstruction of those samples, without reverse-trajectory likelihoods or target audio. On 500 Russian CV3-Eval voice-cloning prompts, joint alignment reduces word error rate from 3.18% to 2.42% at the reward-selected checkpoint (24.0% relative) and to 2.58% at the final checkpoint (19.0%), while WavLM speaker cosine rises from 0.733 to 0.748 and 0.755. Controlled experiments characterize recognition-identity trade-offs and the effects of corruption count, group composition, and weighting.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Maxim Maslov, Kirill Borodin, Vasilii Kudryavtsev, Nikita Vasiliev, Grach Mkrtchian. 2026-09-29. RAWD-TTS: Ratio-Free Reward Alignment for Discrete-Diffusion Voice Cloning. https://arxiv.org/abs/2609.37028
Cite the original work for its findings. Save a collection to share your selection of sources.