arXiv · 2512.05126
SyncVoice: Simple and Effective Automatic Video Dubbing with Vision-Augmented TTS
Abstract
Automatic video dubbing aims to generate high-fidelity speech that is temporally aligned with visual content. However, existing methods still suffer from limited speech naturalness, insufficient audio-visual synchronization, and poor scalability beyond monolingual settings. To address these challenges, we propose SyncVoice, a simple and effective dubbing framework that lightly integrates a Text-Visual Fusion Module into a pretrained text-to-speech (TTS) system. This module aligns visual features with linguistic representations, enabling temporally synchronized speech synthesis without complex architectural redesign. Experiments on the LRS3 dataset show that SyncVoice achieves state-of-the-art performance in zero-shot dubbing. Further training on a large-scale bilingual audio-visual dataset improves vocal fidelity while preserving synchronization, yielding a single unified model for both Chinese and English dubbing.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kaidi Wang, Yi He, Wenhao Guan, Weijie Wu, Peijie Chen, Hongwu Ding, Xiong Zhang, Di Wu, Meng Meng, Jian Luan, Lin Li, Qingyang Hong. 2026-09-15. SyncVoice: Simple and Effective Automatic Video Dubbing with Vision-Augmented TTS. https://arxiv.org/abs/2512.05126
Cite the original work for its findings. Save a collection to share your selection of sources.