arXiv · 2512.21653
Decoder-Side Semantic Conditioning for Low-Bitrate Neural Speech Compression
Abstract
Speech codecs are usually optimized for waveform fidelity, allocating bits to acoustic detail that can be inferred from linguistic structure. This leads to inefficient compression and degraded recognition performance. We propose SemDAC, a semantic-aware neural speech codec that adds hierarchical semantic conditioning to residual vector quantization (RVQ). The first RVQ quantizer is distilled from HuBERT features to produce semantic tokens capturing phonetic content, while later quantizers encode residual acoustics. The decoder is conditioned on semantic tokens via feature-wise linear modulation (FiLM), steering reconstruction toward information not explained by semantic abstraction. At 0.95 kbps, SemDAC matches or surpasses a 2.5 kbps DAC baseline on PESQ, STOI, SI-SNR, and Whisper WER, with comparable ViSQOL, and achieves higher subjective MOS than higher-bitrate DAC baselines. Results show that explicit semantic conditioning, rather than token disentanglement or increased model size alone, improves compression efficiency and recognition robustness.
Explore related subjects
Keep this discovery
Liuyang Bai, Weiyi Lu, Li Guo. 2025-12-25. Decoder-Side Semantic Conditioning for Low-Bitrate Neural Speech Compression. https://arxiv.org/abs/2512.21653
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.