arXiv · 2606.07240
KIT's Submission to Cross-Lingual Voice Cloning in IWSLT 2026
Abstract
Cross-lingual voice cloning aims to generate speech in a target language while preserving speaker identity from a source-language reference. This task is central to speech translation and is the focus of the IWSLT 2026 Cross-Lingual Voice Cloning track. A key challenge is maintaining intelligibility and naturalness in the presence of accent variation and domain-specific vocabulary. We build on a multilingual text-to-speech model, FishAudio-S2-Pro, and introduce language tag prompting to improve language control and reduce accent leakage. We further apply reinforcement learning (RL) fine-tuning for task adaptation and observe improvements in intelligibility. Finally, we propose a reference-conditioned lexical matching method that improves pronunciation of domain-specific terms when lexical overlap is present. Results show that language prompting provides the largest gains, while lexical matching yields consistent improvements on matched subsets.
Explore related subjects
Keep this discovery
Seymanur Akti, Alexander Waibel. 2026-06-05. KIT's Submission to Cross-Lingual Voice Cloning in IWSLT 2026. https://arxiv.org/abs/2606.07240
Cite the original work for its findings. Save a collection to share your selection of sources.