arXiv · 2604.26136
One Voice, Many Tongues: Cross-Lingual Voice Cloning for Scientific Speech
Abstract
Preserving a speaker's voice identity while generating speech in a different language remains a fundamental challenge in spoken language technology, particularly in specialized domains such as scientific communication. In this paper, we address this challenge through our system submission to the International Conference on Spoken Language Translation (IWSLT 2026), the Cross-Lingual Voice Cloning shared task. First, we evaluate several state-of-the-art voice cloning models for cross-lingual speech generation of scientific texts in Arabic, Chinese, and French. Then, we build voice cloning systems based on the OmniVoice foundation model. We employ data augmentation via multi-model ensemble distillation from the ACL 60/60 corpus. We investigate the effect of using this synthetic data for fine-tuning, demonstrating improvements in intelligibility (WER & CER) and speaker similarity (SIM), with gains varying across languages.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Amanuel Gizachew Abebe, Yasmin Moslem. 2026-04-28. One Voice, Many Tongues: Cross-Lingual Voice Cloning for Scientific Speech. https://arxiv.org/abs/2604.26136
Cite the original work for its findings. Save a collection to share your selection of sources.