arXiv · 2603.11578
Streaming Translation and Transcription Through Speech-to-Text Causal Alignment
Abstract
Simultaneous machine translation (SiMT) has traditionally relied on offline machine translation models coupled with human-engineered heuristics or learned policies. We propose Hikari, a policy-free, end-to-end model for simultaneous speech-to-text translation and streaming transcription. We also introduce Decoder Time Dilation, a mechanism that counteracts the overrepresentation of WAIT tokens in training. We present a supervised fine-tuning strategy that trains the model to recover from delays, significantly improving the quality-latency trade-off. Despite its modest size, Hikari delivers competitive translation quality at consistently low latency, comparing favorably with published IWSLT 2026 submissions up to 38x larger and with proprietary API systems across en-ja, en-de, and en-ru. We release our model weights and code to facilitate further research.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Roman Koshkin, Jeon Haesung, Lianbo Liu, Hao Shi, Mengjie Zhao, Yusuke Fujita, Yui Sudo. 2026-03-12. Streaming Translation and Transcription Through Speech-to-Text Causal Alignment. https://arxiv.org/abs/2603.11578
Cite the original work for its findings. Save a collection to share your selection of sources.