arXiv · 2609.37116
Multichannel Audio Quality Assessment: Extending Pretrained Perceptual Models to Spatial Audio
Abstract
Accurate perceptual quality assessment is essential for evaluating and optimizing spatial audio, where perceived quality depends on both signal fidelity and inter-channel spatial relationships. However, subjective evaluation is costly, while existing perceptual models are often trained for limited channel configurations and cannot be directly applied to higher-channel-count audio. This raises the question: how can pretrained perceptual knowledge be effectively reused for multichannel spatial audio? Using 5.1-channel audio, we study four levels of multichannel integration: signal, prediction, latent, and feature and propose two learned approaches: latent-level aggregation of spatial-group representations and the feature-level Feature-Band Group Attention (FGAtt), which adaptively fuses spatial groups at the feature level before perceptual processing. Across five 5.1-channel test sets, FGAtt achieves the strongest over- all performance, demonstrating the effectiveness of feature-level adaptation for reusing pretrained perceptual knowledge
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Gouthaman KV, Shiv Gehlot, Vishnu Raj, Lars Villemoes, Arijit Biswas. 2026-09-29. Multichannel Audio Quality Assessment: Extending Pretrained Perceptual Models to Spatial Audio. https://arxiv.org/abs/2609.37116
Cite the original work for its findings. Save a collection to share your selection of sources.