Search arXiv⌕ Search

arXiv · 2609.31041

Room Impulse Response Embeddings for Speech Enhancement in Noisy and Reverberant Environments

Abstract

We propose a self-supervised approach for learning room impulse response (RIR) representations from single-channel noisy-reverberant speech. It consists of first training on reverberant data, then on noisy-reverberant data, and finally with a teacher-student approach, where the student learns to replicate the teacher's embeddings when given a noisy version of the reverberant input. We assess their representational capabilities by estimating acoustic room parameters from them. Conditioning a discriminative speech enhancement model on the derived embeddings yields consistent gains across all evaluated metrics, including downstream word error rate, for both reverberant and noisy-reverberant speech.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Adrian Meise, Reinhold Haeb-Umbach. 2026-09-25. Room Impulse Response Embeddings for Speech Enhancement in Noisy and Reverberant Environments. https://arxiv.org/abs/2609.31041

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

HuPER: A Human-Inspired Framework for Phonetic Perception

We propose HuPER, a human-inspired framework that models phonetic perception as adaptive inference over acoustic-phonetics evidence and linguistic knowledge. With only 100 hours of training data, HuPER achieves state-of-the-art phonetic error rates on five English benchmarks and strong zero-shot transfer to 95 unseen languages. HuPER is also the first framework to enable adaptive, multi-path phonetic perception under diverse acoustic conditions. All training data, models, and code are open-sourced. Code and demo avaliable at https://github.com/Berkeley-Speech-Group/HuPER.

eess.AS↗

SHroom: A Python Framework for Ambisonics Room Acoustics Simulation and Binaural Rendering

Spatial audio research for virtual and augmented reality, teleconferencing and hearing devices often represents sound fields in the Spherical Harmonics (SH) domain, known as Ambisonics. A typical study simulates a room, renders what a listener or a microphone array would capture in it, and processes those signals in the SH domain. We present SHroom (Spherical Harmonics ROOM), an open-source Python library that performs this whole workflow in one package, from room simulation to binaural rendering, head rotation, microphone-array simulation and Ambisonics encoding. Existing tools cover either the room simulation or the downstream processing, so researchers bridge them with ad-hoc code that is hard to reproduce and compare. In SHroom, every step operates on one shared signal type through one processing interface, so a simulated room flows through the complete chain without format conversions. Built on the image-source engine of `pyroomacoustics`, SHroom reproduces the Ambisonic Room Impulse Response (ARIR) of its spherical-harmonic receivers while computing it about 3x faster for SH orders 4 to 12. SHroom is available at 'https://github.com/Yhonatangayer/shroom' and installable via `pip install pyshroom`.

eess.AS↗

Exploring Second-Order Pattern Recognition in Speaker Recognition

In traditional pattern recognition tasks, neural networks are trained to recognise human-defined patterns (e.g. audio categories) in model inputs (e.g. audio). Meanwhile, some Explainable AI (XAI) methods explain latent patterns characterising the network's recognition of inputs as human-defined patterns; this work calls these latent patterns second-order patterns and proposes to discover them. Accordingly, we apply a hierarchical clustering algorithm to analyse whether our speaker recognition network's representations learned from known utterances naturally form hierarchical clusters. Each resulting cluster is a second-order pattern that characterises a context in our network's recognition of the known utterances as speaker identities. All discovered second-order patterns are then interpreted using the Hierarchical Cluster-Class Matching (HCCM) method. Moreover, we propose a new task, second-order pattern recognition, to identify which of the discovered second-order patterns characterising the recognition of known utterances also apply to unseen utterances, thereby characterising the recognition of unseen utterances. Accordingly, we design the Hierarchical Cluster Navigation and Assignment (HCNA) method. HCNA recognises a second-order pattern as applying to an unseen utterance when the utterance's network representation lies within the extrapolation space of the cluster regarded as that second-order pattern. Experimental results demonstrate that the extrapolation space introduced in HCNA substantially improves task performance.

eess.AS↗