arXiv · 2610.04235
Learning to Watermark Speech Synthesis Against Model-Driven Reconstruction
Abstract
Modern TTS systems increasingly generate synthetic speech at scale for diverse users. This setting calls for content-level provenance that can verify the origin of released speech and attribute it to the requesting user, which generative watermarking can support by embedding multi-bit identifiers directly into synthesized speech. Once released, however, speech may undergo heterogeneous learned transformations during distribution and editing, with reconstruction objectives that can preserve speech utility while affecting watermark recoverability differently. We find that no single watermark carrier remains consistently reliable across reconstruction models, as its survival depends jointly on the embedded structure, reconstruction mechanism, and observation representation. To this end, we propose Thrive, a multi-bit generative speech watermarking framework for modern autoregressive TTS, covering both discrete-token and continuous-representation generation under reconstruction attacks. Specifically, Rise synchronizes watermark injection into intermediate representations with its continued integration into subsequent generation, while Care combines waveform and spectral experts using bit-wise reliability selection. Experiments on both autoregressive paradigms show that Thrive preserves synthesis fidelity, achieves 87.6% average recovery accuracy under reconstruction attacks, and supports source attribution over candidate sets of up to 10,000 identities.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Weizhi Liu, Yue Li, Hui Tian, Zhaoxia Yin. 2026-10-03. Learning to Watermark Speech Synthesis Against Model-Driven Reconstruction. https://arxiv.org/abs/2610.04235
Cite the original work for its findings. Save a collection to share your selection of sources.