Search arXiv⌕ Search

arXiv · 2609.35672

Retrieving Individual Stems from Music Mixtures with Slot Embeddings

Abstract

Music producers search libraries of isolated instrument recordings, called stems, for sounds resembling parts of an existing song. Neural retrieval systems address this by mapping audio to embeddings and ranking library stems by their similarity to the query. The leading method, Contrastive Instrument Retrieval (CIR), encodes the mixture as a single embedding, but it works best when a user specifies the target's instrument family. We introduce Stembed, which encodes a mixture as several slot embeddings representing candidate stems. During training, we construct mixtures from stems of the same song and match their slot embeddings to those of the isolated stems. The slot embeddings from mixtures inherit the stem identities of their assigned solo embedding, enabling a contrastive loss. On mixtures from held out MoisesDB artists, Stembed outperforms a CIR-style baseline when both search the full stem library. Even when predicting the stem count itself without family labels, Stembed exceeds the baseline's family-filtered R@1. Our website demonstrates how users can select a slot by inspecting the tags of its retrieved stems.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

David Braun, Junyi Fan, Pranay Manocha, Donald S. Williamson, Adam Finkelstein. 2026-09-28. Retrieving Individual Stems from Music Mixtures with Slot Embeddings. https://arxiv.org/abs/2609.35672

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment

Building reliable speech systems often requires combining multiple modalities, like audio and visual cues. While such multimodal solutions frequently lead to improvements in performance and may even be critical in certain cases, they come with several constraints such as increased sensory requirements, computational cost, and modality synchronization, to mention a few. These challenges constrain the direct uses of these multimodal solutions in real-world applications. In this work, we develop approaches where the learning happens with all available modalities but the deployment or inference is done with just one or reduced modalities. To do so, we propose a Multimodal Training and Unimodal Deployment (MUTUD) framework which includes a Temporally Aligned Modality feature Estimation (TAME) module that can estimate information from missing modality using modalities present during inference. This innovative approach facilitates the integration of information across different modalities, enhancing the overall inference process by leveraging the strengths of each modality to compensate for the absence of certain modalities during inference. We apply MUTUD to various audiovisual speech tasks and show that it can reduce the performance gap between the multimodal and corresponding unimodal models to a considerable extent. MUTUD can achieve this while reducing the model size and compute compared to multimodal models, in some cases by almost 80%.

cs.SD↗

Beyond Acoustic Prefixes: Persistent Access to Serialized Acoustic Memory for LLM-Based Multi-Talker Speech Recognition

Large language models (LLMs) provide strong linguistic priors for serialized output training (SOT), yet LLM-based multi-talker ASR degrades substantially as the number of overlapping talkers increases. Conventional systems expose acoustic evidence primarily through an initial projected mixture prefix, requiring the decoder to preserve and recover talker-relevant information indirectly throughout autoregressive generation. We first examine whether this limitation can be resolved by enriching the static prefix using discrete connectionist temporal classification (CTC) tokens, hybrid token--acoustic prompts, and continuous talker-specific representations. The results suggest that acoustic content alone does not fully address the conditioning bottleneck. We therefore extend onset-based serialization from the output target to the acoustic-conditioning pathway and introduce persistent decoder-side access to SOT-aligned serialized acoustic memory. Talker-specific representations are organized in utterance-onset order and retained as external acoustic memory, while the conventional mixture prefix provides a complementary global-conditioning path. LLM layers query this memory throughout generation through gated residual cross-attention. We further introduce a second adaptation stage that jointly applies low-rank updates to the acoustic-retrieval pathway and selected LLM self-attention projections. Experiments on LibriMix show consistent improvements over static-prefix prompting. These results indicate that effective LLM-based multi-talker ASR depends not only on providing richer acoustic representations, but also on maintaining persistent access to acoustic evidence structured according to the serialized output.

cs.SD↗

AdaptDuplex: from static to adaptive full-duplex spoken dialogue

Full-duplex spoken dialogue requires simultaneous listening and speaking at sub-second latency, under conversational timing and cognitive demands that change moment to moment. Yet current models mostly impose static operating points, lacking a systematic mechanism for adaptive decisions. We present AdaptDuplex, which extends Qwen3-Omni with such a mechanism, co-designed across three layers. A compact token-level protocol represents every window as a canonical sequence, trains dual-stream alignment through a bounded text lead over speech, and exposes every behavioral decision as an explicit token for training-free runtime control via logits bias. Adaptive mechanisms dynamically predict among discrete window durations and augment direct response as needed with non-blocking cognitive consolidation and multi-flight external reasoning. A progressive pipeline introduces these behaviors through a three-stage Thinker curriculum, then Talker-only and joint SFT, with GRPO as a preliminary increment. On Full-Duplex-Bench v1 and v1.5, AdaptDuplex outperforms DuplexOmni on 18 of 21 comparable turn-taking, overlap-behavior, and timing metrics and MiniCPM-o 4.5 on 19 of 22, with gains in both interaction decisions and response timing. On human-recorded HumDial-FDBench, it attains the top Final score (69.6) of the compared duplex models.

cs.SD↗