arXiv · 2606.22766
READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations
Abstract
Audio Description aims to generate concise narrations of essential visual content in audio-visual media for blind and low-vision audiences. Existing methods either rely on prompting off-the-shelf multimodal models, which often mismatch AD style, or partially optimize training-based systems with next-token prediction, which under-explores model capacity and biases generation toward generic expressions. We present READ, the first reinforcement-learning (RL) framework for training-based AD generation. READ formulates AD as sequence-level optimization with reference-matching, length, and format rewards, and further introduces a dedicated coherence reward under context-aware supervision to promote narratively coherent descriptions. Experiments on MAD-Eval, CMD-AD, and TV-AD show that READ substantially outperforms prior methods across diverse evaluation metrics. Our results highlight RL as a promising paradigm for accurate and coherent AD generation. Our codes, models, and benchmark results will be publicly available.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Bo Fang, Xinyao Zhang, Yuxin Song, Hui Zhang, Hang Zhou, Antoni B. Chan. 2026-06-22. READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations. https://arxiv.org/abs/2606.22766
Cite the original work for its findings. Save a collection to share your selection of sources.