arXiv · 2605.17225
Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities
Abstract
Robust selective auditory attention under multilingual interference is critical for reliable LALM deployment. We introduce MUSI, a cocktail party-inspired controlled diagnostic evaluation of source-grounded spoken-language understanding and reasoning. Each item pairs an English target dialogue with a plausible distractor in English, Spanish, Korean, or Chinese, and evaluates models under (1) single-stream, (2) separation-based, and (3) end-to-end cocktail party settings across controlled SNRs. Across four open-weight and two closed-source LALMs, we find model-dependent language-conditioned variation and heightened vulnerability at adverse SNRs. Errors are dominated by distractor-grounded source confusion, while separation reduces acoustic overlap but often leaves source attribution unresolved. These findings highlight selective auditory attention as an primary capability for reliable LALMs deployment.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Heejoon Koo. 2026-09-17. Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities. https://arxiv.org/abs/2605.17225
Cite the original work for its findings. Save a collection to share your selection of sources.