arXiv · 2609.30540
Don't CLAP: Are Music-Text Models Bag-of-Words?
Abstract
Text-to-music systems are assessed on audio quality and on how faithfully the music follows its prompt, and the CLAP score, the cosine similarity between a music-text model's audio and text embeddings, is the standard objective metric of faithfulness. We ask how accurately that score reflects the text: when an attribute is linked to an instrument (e.g., distorted guitar), does the text embedding capture that binding? To find out, we introduce an attribute swap perturbation: the caption of a real recording is edited by exchanging exactly one property, timbre, lead versus accompaniment, or order of first appearance, between two instruments. We then test four contrastive music-text models and one large audio-language model on whether the audio scores higher against the original caption than against the perturbed one. No contrastive model distinguishes the two captions reliably. The audio-language model does better, but further experiments show that its advantage rests largely on audio-agnostic language priors. Our results thus provide compelling evidence that the CLAP score and related metrics do not capture fine-grained musical meaning or attribute bindings; their representation is closer to a bag-of-words that leaves them insensitive to meaning-changing perturbations of the caption.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yuan-Chiao Cheng, Alexander Lerch. 2026-09-24. Don't CLAP: Are Music-Text Models Bag-of-Words?. https://arxiv.org/abs/2609.30540
Cite the original work for its findings. Save a collection to share your selection of sources.