arXiv · 2605.20946
Thinking-while-speaking: A Controlled, Interleaved Reasoning Method for Real-Time Speech Generation
Abstract
The thinking-while-speaking paradigm aims to make AI communication more human. A key challenge is maintaining fluent speech while performing deep reasoning. Our method, InterRS, tackles this by inserting reasoning steps only during natural speech generation. This requires high-quality data where reasoning and speech are precisely aligned, and the length ratio are under controlled. We introduce a novel pipeline to generate such seamlessly interleaved audio data. To train our model, we combine interleaved SFT with refined data and reinforcement learning with two new rewards: a TA-Balance Reward to manage timing and thinking-answer ratio, and a Linguistic Quality Reward to refine expression. Experiments show our approach achieves 13% better performance on mathmatical and logic benchmarks while generating instant response like a spoken-language instruct model which outputs fast CoT response. Furthermore, our method generates more natural and fluent answers than prior methods.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xuan Du, Qiangyu Yan, Wenshuo Li, Borui Jiang, Changming Xiao, Han Shu, Xinghao Chen. 2026-05-20. Thinking-while-speaking: A Controlled, Interleaved Reasoning Method for Real-Time Speech Generation. https://arxiv.org/abs/2605.20946
Cite the original work for its findings. Save a collection to share your selection of sources.