arXiv · 2605.20946
Thinking-while-speaking: A Controlled, Interleaved Reasoning Method for Real-Time Speech Generation
Abstract
The thinking-while-speaking paradigm aims to make AI communication more human. A key challenge is maintaining fluent speech while performing deep reasoning. Our method, InterRS, tackles this by inserting reasoning steps only during natural speech generation. This requires high-quality data where reasoning and speech are precisely aligned, and the length ratio are under controlled. We introduce a novel pipeline to generate such seamlessly interleaved audio data. To train our model, we combine interleaved SFT with refined data and reinforcement learning with two new rewards: a TA-Balance Reward to manage timing and thinking-answer ratio, and a Linguistic Quality Reward to refine expression. Experiments show our approach achieves 13% better performance on mathmatical and logic benchmarks while generating instant response like a spoken-language instruct model which outputs fast CoT response. Furthermore, our method generates more natural and fluent answers than prior methods.
Explore related subjects
Keep this discovery
Xuan Du, Qiangyu Yan, Wenshuo Li, Borui Jiang, Changming Xiao, Han Shu, Xinghao Chen. 2026-05-20. Thinking-while-speaking: A Controlled, Interleaved Reasoning Method for Real-Time Speech Generation. https://arxiv.org/abs/2605.20946
Cite the original work for its findings. Save a collection to share your selection of sources.