arXiv · 2609.37611
Selective Lookahead for Attention-Based Streaming ASR
Abstract
End-to-end attention-based speech recognition is accurate offline but hard to stream: outputs can depend on future audio, and a little future context per layer makes the lookahead grow with the number of layers. We address this with two mechanisms. A bounded-lookahead chunk encoder caps every chunk's future receptive field at a constant number of chunks, independent of the number of layers, via one age-selection rule shared by self-attention and the depthwise convolution. On this encoder, dynamic future-chunk decoding lets a per-token trigger commit a token or wait and re-decode it; we propose a learned trigger as the general mechanism, with a simple confidence threshold as an effective fallback. On full LibriSpeech test-clean the dynamic system matches the best static-lookahead accuracy (6.5%) at a median latency of 306 ms versus 860 ms for one-chunk static lookahead, and a wait budget bounds the deferral tail below the static baseline's 90th percentile at 0.1 points more WER.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yichen Jia, Bastiaan Tamm, Hugo Van hamme. 2026-09-29. Selective Lookahead for Attention-Based Streaming ASR. https://arxiv.org/abs/2609.37611
Cite the original work for its findings. Save a collection to share your selection of sources.