arXiv · 2605.04749
Spatial-Magnifier: Spatial upsampling for multichannel speech enhancement
Abstract
While the spatial directivity of multichannel speech enhancement algorithms improves with the number of microphones, fitting large capture arrays into real-world edge devices is typically limited by physical constraints. To overcome this limitation, we propose Spatial-Magnifier, a neural network designed to generate virtual microphone (VM) signals from a limited set of real microphone (RM) measurements. Moreover, we introduce the Spatial Audio Representation Learning (SARL) framework, which leverages estimated VM signals and features to condition a downstream speech enhancement system. Experimental results demonstrate that the proposed framework outperforms existing spatial upsampling baselines across various speech extraction systems, including end-to-end multichannel speech enhancement and neural beamforming. The proposed method nearly recovers the oracle performance achieved when all microphones are available.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Dongheon Lee, Ashutosh Pandey, Sanjeel Parekh, Daniel Wong, Jacob Donley, Buye Xu, Juan Azcarreta. 2026-05-06. Spatial-Magnifier: Spatial upsampling for multichannel speech enhancement. https://arxiv.org/abs/2605.04749
Cite the original work for its findings. Save a collection to share your selection of sources.