Search arXiv⌕ Search

arXiv · 2609.29866

Beyond Model Size: Redesigning LiSenNet for embedded speech enhancement

Abstract

Deploying real-time speech enhancement on resource-constrained devices requires meeting strict latency, memory, and energy constraints. Microcontroller NPUs can accelerate neural inference under these constraints, but only through a restricted set of operators in static, integer-quantized graphs. Recent speech-enhancement networks have reduced parameter counts and MACs to levels nominally suitable for microcontrollers, but their operators and execution patterns often remain incompatible with restricted NPUs. We address this gap by redesigning LiSenNet, a 37k parameter sub-band dual-path model, for the STM32N6570-DK Neural-ART accelerator. We replace its recurrent bottleneck with convolutional frequency and temporal mixers, reformulate unsupported operations as static int8-compatible primitives, and use bounded decoder activations to preserve quality after quantization. On VoiceBank-DEMAND, the final NPU-compatible model matches or exceeds the recurrent LiSenNet baseline, reaching PESQ 3.08 versus 3.01 in FP32 and 3.01 versus 2.93 in int8. Deployed on a microcontroller, it processes each 16 ms input hop in 4.83 ms, corresponding to a real-time factor of 0.30. Stateless receptive-field recomputation is an order of magnitude slower at the same frame rate despite higher accelerator utilization. These results show that parameter count and operator compatibility, quantization range, and persistent streaming state must be co-designed to achieve efficient real-time speech enhancement on restricted NPUs.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Clément Laroche, Rasmus Kongsgaard Olsson. 2026-09-24. Beyond Model Size: Redesigning LiSenNet for embedded speech enhancement. https://arxiv.org/abs/2609.29866

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

HuPER: A Human-Inspired Framework for Phonetic Perception

We propose HuPER, a human-inspired framework that models phonetic perception as adaptive inference over acoustic-phonetics evidence and linguistic knowledge. With only 100 hours of training data, HuPER achieves state-of-the-art phonetic error rates on five English benchmarks and strong zero-shot transfer to 95 unseen languages. HuPER is also the first framework to enable adaptive, multi-path phonetic perception under diverse acoustic conditions. All training data, models, and code are open-sourced. Code and demo avaliable at https://github.com/Berkeley-Speech-Group/HuPER.

eess.AS↗

SHroom: A Python Framework for Ambisonics Room Acoustics Simulation and Binaural Rendering

Spatial audio research for virtual and augmented reality, teleconferencing and hearing devices often represents sound fields in the Spherical Harmonics (SH) domain, known as Ambisonics. A typical study simulates a room, renders what a listener or a microphone array would capture in it, and processes those signals in the SH domain. We present SHroom (Spherical Harmonics ROOM), an open-source Python library that performs this whole workflow in one package, from room simulation to binaural rendering, head rotation, microphone-array simulation and Ambisonics encoding. Existing tools cover either the room simulation or the downstream processing, so researchers bridge them with ad-hoc code that is hard to reproduce and compare. In SHroom, every step operates on one shared signal type through one processing interface, so a simulated room flows through the complete chain without format conversions. Built on the image-source engine of `pyroomacoustics`, SHroom reproduces the Ambisonic Room Impulse Response (ARIR) of its spherical-harmonic receivers while computing it about 3x faster for SH orders 4 to 12. SHroom is available at 'https://github.com/Yhonatangayer/shroom' and installable via `pip install pyshroom`.

eess.AS↗

Exploring Second-Order Pattern Recognition in Speaker Recognition

In traditional pattern recognition tasks, neural networks are trained to recognise human-defined patterns (e.g. audio categories) in model inputs (e.g. audio). Meanwhile, some Explainable AI (XAI) methods explain latent patterns characterising the network's recognition of inputs as human-defined patterns; this work calls these latent patterns second-order patterns and proposes to discover them. Accordingly, we apply a hierarchical clustering algorithm to analyse whether our speaker recognition network's representations learned from known utterances naturally form hierarchical clusters. Each resulting cluster is a second-order pattern that characterises a context in our network's recognition of the known utterances as speaker identities. All discovered second-order patterns are then interpreted using the Hierarchical Cluster-Class Matching (HCCM) method. Moreover, we propose a new task, second-order pattern recognition, to identify which of the discovered second-order patterns characterising the recognition of known utterances also apply to unseen utterances, thereby characterising the recognition of unseen utterances. Accordingly, we design the Hierarchical Cluster Navigation and Assignment (HCNA) method. HCNA recognises a second-order pattern as applying to an unseen utterance when the utterance's network representation lies within the extrapolation space of the cluster regarded as that second-order pattern. Experimental results demonstrate that the extrapolation space introduced in HCNA substantially improves task performance.

eess.AS↗