arXiv · 2609.33443
Context Spanning: A Communication Framework for Full-Duplex Speech Models and External LLM Backends
Abstract
Full-duplex spoken dialogue models can listen and speak simultaneously like the real-time dynamics of human conversation. For natural dialogue, the ability to search for external information in real-time is also an important capability. Many models remain trapped in parametric knowledge, leaving them unable to access real-time information and tool execution. Furthermore, even when Large Language Models (LLM) retrieve information, many duplex speech models process it within a compressed latent space rather than in its raw text form, which can lead to information loss from compression. To address this issue, we propose Context Spanning, a framework for information injection between a full-duplex speech model and an external LLM backend via real-time chunked prefill. The injected frame is encoded in a single forward pass inside the real-time frame budget. It feeds the retrieved information to the speech model as-is, enabling it to reason over the information independently and generate responses. With this approach, our model achieves high performance on Full-Duplex benchmarks and strong results on Question Answering tasks, demonstrating its conversation potential. Context Spanning shows that external information can be injected directly into a duplex speech model, introducing a new simple and powerful mechanism for duplex systems.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Seonghyeon Go, Yongwoo Kim, Hyeonjin Cha, Jaeho Shin. 2026-09-27. Context Spanning: A Communication Framework for Full-Duplex Speech Models and External LLM Backends. https://arxiv.org/abs/2609.33443
Cite the original work for its findings. Save a collection to share your selection of sources.