arXiv · 2609.26255
SceneTTS-Bench: A Benchmark for Scene-Level TTS in Drama Dubbing
Abstract
Text-to-speech systems are increasingly used for drama dubbing, yet evaluation protocols remain sentence-level, leaving critical scene-level behaviors insufficiently measured. We present SceneTTS-Bench, a benchmark that evaluates TTS along three dimensions: timbre consistency across character turns, emotional expressiveness on high-tension utterances, and rhythm coherence under segmented long-form synthesis. The corpus comprises real-world and generated drama scripts totaling 160 bilingual scenes and approximately 10,300 utterances, with real-world scripts serving as the primary source (100 scenes) and generated scripts as a supplementary source (60 scenes), demonstrating the framework's extensibility through synthetic data augmentation. A backend-agnostic Canonical Intermediate Representation ensures fair cross-system comparison. Three automatic pipelines produce per-utterance diagnostics: Speaker Consistency Score for timbre-drift detection, Under-Acting Ratio for under-acting identification, and Rate Discontinuity Ratio for rate-discontinuity quantification. Experiments on four TTS systems confirm that each system exhibits distinct weaknesses and that scene-level rankings diverge substantially from sentence-level metrics. Benchmark resources are publicly available at https://piedpiperg.github.io/scenetts-bench/ .
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yizhong Geng, Yanliang Li, Jinghan Yang, Tianhan Jiang, Yingming Gao, Ya Li. 2026-08-13. SceneTTS-Bench: A Benchmark for Scene-Level TTS in Drama Dubbing. https://arxiv.org/abs/2609.26255
Cite the original work for its findings. Save a collection to share your selection of sources.