Search arXiv⌕ Search

arXiv · 2609.33554

An Efficient Parametric Codec for Low-Bitrate First-Order Ambisonics

Abstract

Driven by the rapid growth of immersive teleconferencing and generative spatial audio, efficient low-bitrate coding of first-order Ambisonics (FOA) has become increasingly important. In this work, we develop a lightweight parametric FOA codec that retains the standard Directional Audio Coding analysis and synthesis while redesigning the spatial metadata quantization scheme. Rather than quantizing direction-of-arrival (DOA) and diffuseness independently, we combine them into a 3-D directivity vector and jointly quantize these vectors across frequency bands via residual vector quantization (RVQ). The RVQ codebooks are optimized within minutes via stage-wise k-means without backpropagation, yielding a constant-bitrate (CBR) representation that decouples metadata rate from the number of frequency bands. Evaluations show that our approach outperforms low-bitrate perceptual codecs in FOA reconstruction, remains competitive with neural codecs on the downstream Sound Event Localization and Detection task, and maintains robust performance when paired with different external monaural codecs. Given its lightweight training, strong performance, and CBR design, we consider the proposed method to be a favorable and reproducible baseline for future FOA codec research.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wei-Ting Lai, Amy Bastine, Lachlan Birnie, Thushara D. Abhayapala, Prasanga N. Samarasinghe. 2026-09-27. An Efficient Parametric Codec for Low-Bitrate First-Order Ambisonics. https://arxiv.org/abs/2609.33554

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

NVV-Locator: From Transcript Tags to Acoustic Boundaries for Fine-Grained Nonverbal Vocalization Grounding

Human speech includes nonverbal vocalizations (NVVs), such as laughter, sighs, breaths, and coughs, which convey affective and interactional information. Existing approaches typically represent NVVs as transcript-level tags, providing limited supervision for their waveform-time boundaries. We present NVV-Locator for fine-grained NVV temporal grounding. We first unify 26 NVV categories across public resources and construct large-scale timestamp-supervised training data through dual-LLM verification, transcript-guided forced alignment, and energy-based boundary refinement. We further introduce NVV-TimeBench, an expert-refined benchmark with 667 utterances and 1,094 events. NVV-Locator uses a non-autoregressive slot-filling architecture to jointly predict lexical timestamps, NVV categories, and event boundaries. On NVV-TimeBench, it achieves 71.0% Micro F1, 70.2% Macro F1, 80.4% Macro mIoU, and 59.6 ms Macro mMAE, outperforming the evaluated large audio model counterparts. Evaluation on an external corpus further demonstrates the cross-corpus generalization of NVV-Locator.

eess.AS↗

A barrier or a booster? Familiarity effects on Mandarin emotion prosody recognition using AI-powered voice cloning

Emotion prosody perception requires simultaneous processing of acoustic cues and speaker identity. While listeners effortlessly decode natural speech, AI synthetic voices introduce cognitive complexities due to subtle acoustic atypicalities. It remains unclear how these synthetic features interact with a listener's prior social knowledge and memory of a familiar speaker. This study investigated how speech sources (human vs. AI) and speaker familiarity affect emotion recognition accuracy and cognitive load. A within-subject task with Mandarin-speaking adults evaluated behavioral (accuracy, reaction time) and physiological data (heart rate variability). Results showed that human voices yielded significantly higher accuracy and faster processing times than AI voices, while HRV did not significantly differentiate between conditions. These findings show that decoding synthetic speech is gated by top-down social cognition, highlighting limitations in current AI synthesis technologies.

eess.AS↗

Improving Predicted MOS Scores, Not Perceived Quality: Multi-Predictor Test-Time Optimization of Enhanced Speech

Non-intrusive MOS predictors are widely used instead of subjective listening tests to evaluate and rank speech enhancement (SE) systems. If they accurately reflect perceived quality, raising their scores should lead to higher-quality speech. We present the first comprehensive analysis of test-time optimization for the SE task, which directly modifies the enhanced signal to raise the average of multiple MOS predictor scores. On seven systems from the URGENT 2026 challenge, we find that 1)~all the optimized predicted scores increase while reference-based metrics remain nearly unchanged, 2)~a non-optimized predicted score does not increase, and 3)~a MUSHRA listening test shows no improvement in perceived quality. These findings reveal a risk that such optimization can distort evaluations, e.g., biasing comparisons of SE systems regardless of their perceived quality. We believe these findings can inform future evaluation practices: they suggest that predictors used for optimization should not be used for evaluation, and that challenges should keep the predictors used for ranking undisclosed.

eess.AS↗