arXiv · 2609.35345
Probing Large Audio-Language Models for Compositional Understanding of Sounding Actions
Abstract
Large audio-language models (LALMs) excel at understanding and reasoning tasks over atomic sound events, yet their ability to infer higher-level human activities from such fine-grained events remains largely unexamined. Everyday human actions and activities, such as setting a table, cleaning the house, or preparing a breakfast emerge compositionally from temporally distributed sound events, requiring abstraction beyond the event-centric granularity that dominates current training and evaluation paradigms. Our benchmark evaluates a wide set of LALMs under a principled framework that tests how language-based reasoning, grounded in acoustic perception, structures sound abstractions into higher-level understanding. By systematically varying exemplar typicality and distractor similarity, our evaluation exposes \added{that current models do not reliably perform compositional inference from atomic acoustic events to higher-level human activities solely from audio.} All data, taxonomies, and evaluation scripts are publicly available on our companion website: https://alm-sounding-actions.onrender.com/
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Michel Olvera, Paraskevas Stamatiadis, Changhong Wang, Ga{ë}l Richard. 2026-09-28. Probing Large Audio-Language Models for Compositional Understanding of Sounding Actions. https://arxiv.org/abs/2609.35345
Cite the original work for its findings. Save a collection to share your selection of sources.