Search arXiv⌕ Search

arXiv subjects

Mengyao Zhu

Publications and source records attributed to Mengyao Zhu.

7 recordsLinked to original sources

RMS-AQA: A Two-Stage Spatial Audio Question Answering Benchmark for Real-World Domestic Environments

Embodied assistants in domestic environments must infer what happened, where and when it occurred, and how to respond. To address this, we introduce RMS-AQA, a spatial audio question answering (SAQA) benchmark for real-world domestic environments. The benchmark features a two-stage question-answering (QA) format to comprehensively assess the ability of audio-language models (ALMs) to first ground audible sound events and subsequently perform complex spatio-temporal reasoning based on that grounding. To maximize acoustic realism, our dataset combines authentic real-world first-order Ambisonics (FOA) recordings with high-fidelity synthetic data generated using measured room impulse responses (RIRs). Furthermore, we provide a lightweight spatial plug-in that injects FOA-format data into frozen audio-language backbones. Experimental results reveal that the primary challenges stem from concurrent sources, far distance, and sim-to-real domain gap between RIR-synthesized and authentic recordings.

cs.SD↗

Signal-Independent and Signal-Dependent Neural Ambisonic Matrix Encoding for Arbitrary Arrays with Variable Microphone Counts

Recent neural Ambisonic encoders accommodate diverse array geometries, yet many existing neural encoders require a fixed microphone count because the number of microphone channels is embedded in the network architecture. This requirement limits deployment across devices with different microphone configurations and adaptation to changes in available channels. To address this limitation, we investigate Transformer-based matrix encoding for arbitrary microphone arrays with variable microphone counts. This is achieved through shared microphone-wise processing and masked self-attention that models inter-microphone relationships across variable-size arrays. Within this framework, we consider signal-independent (SI) encoding, which predicts encoding matrices from array transfer functions, and introduce a signal-dependent (SD) extension that additionally incorporates the observed microphone signals. Both models are trained on simulated scenes using LibriSpeech sources and extensively evaluated under changes in source type, unseen microphone counts, and increased source counts beyond those used during training. Both SI and SD outperform conventional least-squares (LS) encoding in aggregate reconstruction performance across the evaluated conditions. SD consistently achieves stronger overall performance than SI. These results demonstrate that the proposed framework enables array-agnostic Ambisonic encoding while retaining generalization across microphone counts and acoustic source conditions.

eess.AS↗

GAMF: Learned and Analytical Array Transfer Function Matching for Array-Generic Direction-of-Arrival Estimation

Microphone positional encoding supports cross-array direction-of-arrival (DOA) estimation, but coordinates alone cannot fully describe device shadowing or microphone directivity. We propose a Generalizable ATF Matching Framework (GAMF) for DOA estimation across array geometries and microphone counts, using array transfer functions (ATFs) as acoustic descriptors. The learned branch incorporates ATF embeddings into geometry-conditioned neural estimation to match acoustic observations with candidate directions. The analytical branch performs normalized ATF matching adapted from generalized steered response power. A hybrid configuration combines their scores through adaptive gating. Simulations across array configurations show that both learned and hybrid configurations outperform a representative positional-encoding-based neural baseline, remain competitive with analytical ATF matching in clean, low-reverberation scenes, and substantially improve upon it under stronger noise or reverberation. On eight-microphone LOCATA Task 1 recordings, the hybrid configuration outperforms the evaluated state-of-the-art baselines for three-dimensional DOA estimation, achieving mean errors of $3.76^\circ$ for three-dimensional DOA and $2.87^\circ$ for azimuth.

eess.AS↗

PRIME-ANC: Path-Ratio-Informed Modeling for Efficient Neural Filter Synthesis in Active Noise Control

Changes in listener acoustics require active noise control (ANC) filters to be redesigned for new acoustic paths. We introduce PRIME-ANC, a shared neural synthesizer that learns a bounded, path-dependent logmagnitude correction to a regularized path-ratio base. Minimum-phase reconstruction and truncation produce finite-impulse-response (FIR) filters; training optimizes their noise-control performance. Across ten random training/test splits per dataset, PRIME-ANC achieves average held-out one-third-octave reductions of 18.81 and 17.76 dB over 50 Hz-5 kHz on a ten-path dataset and a public earphone database, respectively. On original earphone measurements, it improves upon the path-ratio base by 7.89 dB. Ablation studies support the contributions of both the analytic base and the path-dependent correction. Given calibrated paths for a held-out listener condition, PRIME-ANC generates a path-specific FIR without iterative optimization. With three Gauss-Newton updates, it reaches 21.43 dB reduction, approaching direct weighted least-squares design while producing lower amplification and root-mean-square control output.

eess.AS↗

The CCF AATC 2025 Speech Restoration Challenge: A Retrospective

Real-world speech communication is rarely affected by a single type of degradation. Instead, it suffers from a complex interplay of acoustic interference, codec compression, and, increasingly, secondary artifacts introduced by upstream enhancement algorithms. To bridge the gap between academic research and these realistic scenarios, we introduced the CCF AATC 2025 Challenge. This challenge targets universal blind speech restoration, requiring a single model to handle three distinct distortion categories: acoustic degradation, codec distortion, and secondary processing artifacts. In this paper, we provide a comprehensive retrospective of the challenge, detailing the dataset construction, task design, and a systematic analysis of the 25 participating systems. We report three key findings that define the current state of the field: (1) Efficiency vs. Scale: Contrary to the trend of massive generative models, top-performing systems demonstrated that lightweight discriminative architectures (<10M parameters) can achieve state-of-the-art performance, balancing restoration quality with deployment constraints. (2) Generative Trade-off: While generative and hybrid models excel in theoretical perceptual metrics, breakdown analysis reveals they suffer from "reconstruction bias" in high-SNR codec tasks and struggle with hallucination in complex secondary artifact scenarios. (3) Metric Gap: Most critically, our rank correlation analysis exposes a strong negative correlation (\r{ho}=-0.8) between widely-used reference-free metrics (e.g., DNSMOS) and human MOS when evaluating hybrid systems. This indicates that current metrics may over-reward artificial spectral smoothness at the expense of perceptual naturalness. This paper aims to serve as a reference for future research in robust speech restoration and calls for the development of next-generation evaluation metrics sensitive to generative artifacts.

cs.SD↗

End-to-End Model for Speech Enhancement by Consistent Spectrogram Masking

Recently, phase processing is attracting increasinginterest in speech enhancement community. Some researchersintegrate phase estimations module into speech enhancementmodels by using complex-valued short-time Fourier transform(STFT) spectrogram based training targets, e.g. Complex RatioMask (cRM) [1]. However, masking on spectrogram would violentits consistency constraints. In this work, we prove that theinconsistent problem enlarges the solution space of the speechenhancement model and causes unintended artifacts. ConsistencySpectrogram Masking (CSM) is proposed to estimate the complexspectrogram of a signal with the consistency constraint in asimple but not trivial way. The experiments comparing ourCSM based end-to-end model with other methods are conductedto confirm that the CSM accelerate the model training andhave significant improvements in speech quality. From ourexperimental results, we assured that our method could enha

cs.SD↗

End-to-End Residual CNN with L-GM Loss Speaker Verification System

We propose an end-to-end speaker verification system based on the neural network and trained by a loss function with less computational complexity. The end-to-end speaker verification system in this paper consists of a ResNet architecture to extract features from utterance, then produces utterance-level speaker embeddings, and train using the large-margin Gaussian Mixture loss function. Influenced by the large-margin and likelihood regularization, large-margin Gaussian Mixture loss function benefits the speaker verification performance. Experimental results demonstrate that the Residual CNN with large-margin Gaussian Mixture loss outperforms DNN-based i-vector baseline by more than 10% improvement in accuracy rate.

cs.SD↗