arXiv · 2610.04333
Factorized Delayed Streams Modeling for LLM-based Streaming ASR
Abstract
Delayed Streams Modeling (DSM) enables LLM-based streaming automatic speech recognition (ASR) by aligning acoustic and text streams on a common timeline. DSM adds the padding token and the word-start token to the LLM vocabulary and predicts them together with normal text tokens using the same softmax. We first show that can be removed while maintaining competitive recognition performance. Based on this result, we propose Factorized DSM (F-DSM), which separates the waiting probability for from the distribution over the original LLM vocabulary. This factorization removes ASR-specific tokens from the text prediction space and allows the large-vocabulary softmax to be skipped on waiting steps. Experiments on the Corpus of Spontaneous Japanese and LibriSpeech show that F-DSM achieves better recognition performance than DSM. It also greatly reduces GPU memory use while maintaining similar training throughput, provides a small inference speed improvement through softmax skipping, and reduces the degradation in text-only perplexity observed with DSM.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Tatsunari Takagi, Kai Washizaki, Atsushi Kojima, Lianbo Liu, Koki Nikaido, Yui Sudo. 2026-10-03. Factorized Delayed Streams Modeling for LLM-based Streaming ASR. https://arxiv.org/abs/2610.04333
Cite the original work for its findings. Save a collection to share your selection of sources.