arXiv · 2506.13396
Bi-directional Context-Enhanced Speech Large Language Models for Multilingual Conversational ASR
Abstract
This paper introduces the integration of language-specific bi-directional context into a speech large language model (SLLM) to improve multilingual continuous conversational automatic speech recognition (ASR). We propose a character-level contextual masking strategy during training, which randomly removes portions of the context to enhance robustness and better emulate the flawed transcriptions that may occur during inference. For decoding, a two-stage pipeline is utilized: initial isolated segment decoding followed by context-aware re-decoding using neighboring hypotheses. Evaluated on the 1500-hour Multilingual Conversational Speech and Language Model (MLC-SLM) corpus covering eleven languages, our method achieves an 18% relative improvement compared to a strong baseline, outperforming even the model trained on 6000 hours of data for the MLC-SLM competition. These results underscore the significant benefit of incorporating contextual information in multilingual continuous conversational ASR.
Explore related subjects
Keep this discovery
Yizhou Peng, Hexin Liu, Eng Siong Chng. 2025-06-16. Bi-directional Context-Enhanced Speech Large Language Models for Multilingual Conversational ASR. https://arxiv.org/abs/2506.13396
Cite the original work for its findings. Save a collection to share your selection of sources.