arXiv · 2609.34902
Dual-Stream Simultaneous Translation via 2D Grid Attention
Abstract
Simultaneous machine translation must generate target tokens before the source input is complete. Existing approaches address this through post-hoc read-write policies, leaving the attention mechanism unaware of bidirectional stream dependencies. We propose a dual-stream attention framework that represents source and target streams as a two-dimensional grid of hidden states and models their interaction through four structurally distinct attention types merged via joint QK Softmax normalization. Two approximations---broadcast and Hadamard---reduce the per-layer complexity from O(X^2Y+XY^2) to O(X^2+Y^2+XY) with provably decaying error. Training uses a self-guided loop: a per-cell loss heatmap drives dynamic-programming path recovery, which generates read/write decision supervision labels without external alignment. An incremental KV cache with anchored rotary position embeddings enables efficient streaming inference. On Chinese-to-English simultaneous translation, the proposed model outperforms the Wait-k baseline by +5.66 BLEURT and +10.36 COMET at comparable latency, and surpasses the non-streaming reference on COMET at a fraction of the response delay.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yu Pu, Wei-Qiang Zhang. 2026-09-28. Dual-Stream Simultaneous Translation via 2D Grid Attention. https://arxiv.org/abs/2609.34902
Cite the original work for its findings. Save a collection to share your selection of sources.