arXiv · 2608.27817
Auditing Generative Audio Calls for Known-Task Audio-LLM Evaluatio
Abstract
Speech and audio LLMs are often evaluated by asking whether a waveform prompt beats an automatic speech recognition (ASR) transcript. For known closed-set tasks, that comparison conflates two factors: access to acoustic evidence and the need to call a generative audio model. We evaluate this distinction as a controlled call-decision problem. For each example, a policy chooses among keeping a transcript label, using encoder evidence from Contrastive Language-Audio Pretraining (CLAP), Audio Spectrogram Transformer (AST), or WavLM, and calling Qwen2-Audio, Qwen2.5-Omni, or MOSS-Audio; the decisive ablation removes all generative actions while keeping the selector and development protocol fixed. On VocalSound, the transcript-only representation reaches 0.296 accuracy, so it is insufficient for this label set. Yet supervised CLAP and WavLM controls reach 0.850 and 0.854 with no generative audio calls. A selector with generative actions reaches 0.925 accuracy using 12.5% calls, compared with 0.921 for the matched no-call selector (paired difference 0.004; 95% CI [-0.025, 0.033], including zero). Agreement and stacking features improve weaker selectors but do not beat the strongest no-call control. For known-task endpoint-value statements, the relevant quantity is the marginal value of the generative call after transcript and encoder evidence have already been used.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mengzhe Geng. 2026-09-12. Auditing Generative Audio Calls for Known-Task Audio-LLM Evaluatio. https://arxiv.org/abs/2608.27817
Cite the original work for its findings. Save a collection to share your selection of sources.