arXiv · 2609.36913
BaLEEN: Biasing with Latent Encoded Entities for Context-Aware ASR
Abstract
Transcribing domain-specific entities and rare proper nouns remains a major challenge in automatic speech recognition (ASR). In this paper, we propose BaLEEN (Biasing with Latent Encoded Entities), a lightweight, hypernetwork-based framework for dynamic contextual adaptation without fine-tuning the underlying ASR model. BaLEEN encodes variable-length contextual keywords using a pretrained language model, compresses them into a fixed sequence of latent vectors via a Perceiver bottleneck, and injects context-dependent bias vectors directly into the intermediate encoder representations of the ASR model. Because both the language model and the backbone ASR model remain entirely frozen during training, BaLEEN operates as a plug-and-play adapter that incurs zero computational overhead at inference time when context biases are precomputed. We evaluate our method on a CTC-based ASR model using a Wikipedia-derived corpus with annotated named entities and synthetic speech. Experimental results demonstrate that BaLEEN reduces keyword miss rate by 8.7% on the test set relative to the unbiased baseline while simultaneously improving overall word error rate by 21% and character error rate by 28%.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chihiro Taguchi, Yotaro Kubo, Rujikorn Charakorn. 2026-09-29. BaLEEN: Biasing with Latent Encoded Entities for Context-Aware ASR. https://arxiv.org/abs/2609.36913
Cite the original work for its findings. Save a collection to share your selection of sources.