Search arXiv⌕ Search

arXiv · 2609.32755

SAGE: Semantic Audio Generative Encoder

Abstract

Audio autoencoders compress waveforms into compact latent representations that serve as the interface between raw audio and downstream models. Current systems navigate a three-way trade-off between reconstruction quality, semantic structure of the latent space, and inference speed, typically favoring one or two of these at the expense of the others. This paper introduces SAGE, Semantic Audio Generative Encoder: a compact variational autoencoder, trained solely on publicly available music, that shapes its latent by distilling embeddings from a pretrained audio-text model. This 105M-parameter model runs at the inference cost of Stable Audio Open and reaches the listening-test quality of SAME-L, an autoencoder 8x larger and 4x slower, while surpassing both on objective perceptual and distributional metrics of reconstruction. Furthermore, it sets the state of the art on all nineteen probing tasks of latent semantics, in domain and out of domain. These results establish SAGE as a lightweight audio autoencoder that strikes the best balance of the three-way trade-off among those we evaluate, combining high reconstruction fidelity, state-of-the-art semantic structure, and fast inference.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Francesco Brigante, Luca Cerovaz, Davide Marincione, Giorgio Strano, Luca Zhou, Emanuele Rodolà, Michele Mancusi. 2026-09-26. SAGE: Semantic Audio Generative Encoder. https://arxiv.org/abs/2609.32755

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment

Building reliable speech systems often requires combining multiple modalities, like audio and visual cues. While such multimodal solutions frequently lead to improvements in performance and may even be critical in certain cases, they come with several constraints such as increased sensory requirements, computational cost, and modality synchronization, to mention a few. These challenges constrain the direct uses of these multimodal solutions in real-world applications. In this work, we develop approaches where the learning happens with all available modalities but the deployment or inference is done with just one or reduced modalities. To do so, we propose a Multimodal Training and Unimodal Deployment (MUTUD) framework which includes a Temporally Aligned Modality feature Estimation (TAME) module that can estimate information from missing modality using modalities present during inference. This innovative approach facilitates the integration of information across different modalities, enhancing the overall inference process by leveraging the strengths of each modality to compensate for the absence of certain modalities during inference. We apply MUTUD to various audiovisual speech tasks and show that it can reduce the performance gap between the multimodal and corresponding unimodal models to a considerable extent. MUTUD can achieve this while reducing the model size and compute compared to multimodal models, in some cases by almost 80%.

cs.SD↗

Beyond Acoustic Prefixes: Persistent Access to Serialized Acoustic Memory for LLM-Based Multi-Talker Speech Recognition

Large language models (LLMs) provide strong linguistic priors for serialized output training (SOT), yet LLM-based multi-talker ASR degrades substantially as the number of overlapping talkers increases. Conventional systems expose acoustic evidence primarily through an initial projected mixture prefix, requiring the decoder to preserve and recover talker-relevant information indirectly throughout autoregressive generation. We first examine whether this limitation can be resolved by enriching the static prefix using discrete connectionist temporal classification (CTC) tokens, hybrid token--acoustic prompts, and continuous talker-specific representations. The results suggest that acoustic content alone does not fully address the conditioning bottleneck. We therefore extend onset-based serialization from the output target to the acoustic-conditioning pathway and introduce persistent decoder-side access to SOT-aligned serialized acoustic memory. Talker-specific representations are organized in utterance-onset order and retained as external acoustic memory, while the conventional mixture prefix provides a complementary global-conditioning path. LLM layers query this memory throughout generation through gated residual cross-attention. We further introduce a second adaptation stage that jointly applies low-rank updates to the acoustic-retrieval pathway and selected LLM self-attention projections. Experiments on LibriMix show consistent improvements over static-prefix prompting. These results indicate that effective LLM-based multi-talker ASR depends not only on providing richer acoustic representations, but also on maintaining persistent access to acoustic evidence structured according to the serialized output.

cs.SD↗

AdaptDuplex: from static to adaptive full-duplex spoken dialogue

Full-duplex spoken dialogue requires simultaneous listening and speaking at sub-second latency, under conversational timing and cognitive demands that change moment to moment. Yet current models mostly impose static operating points, lacking a systematic mechanism for adaptive decisions. We present AdaptDuplex, which extends Qwen3-Omni with such a mechanism, co-designed across three layers. A compact token-level protocol represents every window as a canonical sequence, trains dual-stream alignment through a bounded text lead over speech, and exposes every behavioral decision as an explicit token for training-free runtime control via logits bias. Adaptive mechanisms dynamically predict among discrete window durations and augment direct response as needed with non-blocking cognitive consolidation and multi-flight external reasoning. A progressive pipeline introduces these behaviors through a three-stage Thinker curriculum, then Talker-only and joint SFT, with GRPO as a preliminary increment. On Full-Duplex-Bench v1 and v1.5, AdaptDuplex outperforms DuplexOmni on 18 of 21 comparable turn-taking, overlap-behavior, and timing metrics and MiniCPM-o 4.5 on 19 of 22, with gains in both interaction decisions and response timing. On human-recorded HumDial-FDBench, it attains the top Final score (69.6) of the compared duplex models.

cs.SD↗