arXiv · 2603.09232
How Contrastive Decoding Enhances Large Audio Language Models
Abstract
While Contrastive Decoding (CD) has been proposed to enhance Large Audio Language Models (LALMs), it has not been evaluated at scale, and the underlying mechanisms driving its success remain unclear. This study systematically evaluates four distinct CD strategies across diverse LALM architectures. We identify Audio-Aware Decoding and Audio Contrastive Decoding as the most effective methods. However, their impact varies significantly across models. To explain this variability, we profile the baseline error composition of each model and measure how readily contrastive decoding corrects each error type. Our analysis demonstrates that CD reliably rectifies errors in which models falsely claim an absence of audio or resort to uncertainty-driven guessing, but is relatively poor at correcting flawed reasoning or confident misassertions. Crucially, CD's benefit closely tracks the composition of a model's baseline error profile: when errors caused by audio ignorance or uncertainty-driven guessing constitute only a small fraction of a model's errors, gains are marginal or even negative. A token-level analysis reveals the underlying mechanism: when the amateur's output is dominated by hesitation markers, CD's suppression naturally targets uncertainty-driven errors while having limited effect on confident misassertions.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Tzu-Quan Lin, Wei-Ping Huang, Yi-Cheng Lin, Hung-yi Lee. 2026-09-13. How Contrastive Decoding Enhances Large Audio Language Models. https://arxiv.org/abs/2603.09232
Cite the original work for its findings. Save a collection to share your selection of sources.