arXiv · 2609.31961
Improving Audiovisual Speech Recognition through Synthetic Visual Data Augmentation
Abstract
Audiovisual Speech Recognition (AVSR) is a multimodal approach to speech recognition that incorporates visual information from lip movements to enhance model performance. Despite its advantages, its development remains constrained by the limited availability of labeled audiovisual (AV) datasets. This work explores the use of synthetic visual data as a solution, using an audio-driven talking-head pipeline to generate lip-synchronized visual content from existing audio data. We evaluate the effectiveness of synthetic visual data both as an augmentation strategy and as a standalone training resource, applying our approach to Spanish and Catalan. Our results show that augmenting real AV data with synthetic samples yields relative Word Error Rate (WER) reductions of up to 16.2%, demonstrating the potential of this approach. Moreover, we demonstrate that synthetic data alone can serve as a baseline for AVSR training in languages lacking AV datasets. These findings provide evidence that synthetic visual data can serve as a scalable solution to AVSR data scarcity, enabling broader language coverage.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Pol Buitrago, Pol Gàlvez, Javier Hernando. 2026-09-25. Improving Audiovisual Speech Recognition through Synthetic Visual Data Augmentation. https://arxiv.org/abs/2609.31961
Cite the original work for its findings. Save a collection to share your selection of sources.