Search arXiv⌕ Search

arXiv · 2610.04333

Factorized Delayed Streams Modeling for LLM-based Streaming ASR

Abstract

Delayed Streams Modeling (DSM) enables LLM-based streaming automatic speech recognition (ASR) by aligning acoustic and text streams on a common timeline. DSM adds the padding token and the word-start token to the LLM vocabulary and predicts them together with normal text tokens using the same softmax. We first show that can be removed while maintaining competitive recognition performance. Based on this result, we propose Factorized DSM (F-DSM), which separates the waiting probability for from the distribution over the original LLM vocabulary. This factorization removes ASR-specific tokens from the text prediction space and allows the large-vocabulary softmax to be skipped on waiting steps. Experiments on the Corpus of Spontaneous Japanese and LibriSpeech show that F-DSM achieves better recognition performance than DSM. It also greatly reduces GPU memory use while maintaining similar training throughput, provides a small inference speed improvement through softmax skipping, and reduces the degradation in text-only perplexity observed with DSM.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Tatsunari Takagi, Kai Washizaki, Atsushi Kojima, Lianbo Liu, Koki Nikaido, Yui Sudo. 2026-10-03. Factorized Delayed Streams Modeling for LLM-based Streaming ASR. https://arxiv.org/abs/2610.04333

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

MOV-AAD: A Large-Scale Multimodal Dataset for Auditory Attention Decoding During Moving Conversations

Auditory attention decoding (AAD) is often evaluated on static, simplified speech scenes that poorly match everyday listening. We introduce MOV-AAD, a large-scale dataset for studying auditory attention under moving, naturalistic conversations. MOV-AAD combines 64-channel EEG with synchronized physiological recordings, including eye tracking, respiration, galvanic skin response, heart rate, peripheral oxygen saturation, body temperature, body motion, and photoplethysmography, enabling analysis of cross-modal neural and physiological markers of attention and listening effort. This dataset uses a more ecologically valid listening paradigm with dynamically moving conversational speech sources and behavioral measures of attentional engagement. MOV-AAD supports research on robust AAD in realistic spatial dynamics, multimodal attention modeling, listening effort, and intersubject neural responses, providing a resource for benchmarking selective auditory attention in naturalistic listening.

eess.AS↗

Controllable Accent Normalization via Discrete Diffusion

Existing accent normalization methods do not typically offer control over accent strength, yet many applications-such as language learning and dubbing-require tunable accent retention. We propose DLM-AN, a controllable accent normalization system built on masked discrete diffusion over self-supervised speech tokens. A Common Token Predictor identifies source tokens that likely encode native pronunciation; these tokens are selectively reused to initialize the reverse diffusion process. This provides a simple yet effective mechanism for controlling accent strength: reusing more tokens preserves more of the original accent. DLM-AN further incorporates a flow-matching Duration Ratio Predictor that automatically adjusts the total duration to better match the native rhythm. Experiments on multi-accent English data show that DLM-AN achieves the lowest word error rate among all compared systems while delivering competitive accent reduction and smooth, interpretable accent strength control. The implementation is available at https://github.com/P1ping/DLM-AN

eess.AS↗

FASTDIAR: Frame-level speaker encoder for Streaming Diarization

Real-time conversational agents require speaker diarization that streams and runs on a CPU. Most systems apply an utterance-level speaker encoder to short, heavily overlapping chunks, which wastes computation and leaves the model optimized for the wrong task. We instead turn a state-of-the-art speaker recognition architecture into a causal frame-level encoder that reads the stream once and emits one embedding every 80~ms from a bounded two-second window of past audio, and pair it with online clustering that gates every update on the self-similarity of the stream. Trained only by distillation from an utterance-level teacher on simulated and out-of-domain mixtures, and evaluated with one fixed set of hyperparameters, the system is the most accurate streaming diarizer on low-overlap benchmarks at sub-second latency, degrades far less than cache-based systems as the number of speakers grows, and runs five times faster than real time on a single CPU thread.

eess.AS↗