arXiv · 2608.27360
Your Voice Cloning System is Secretly a Voice Anonymizer
Abstract
Speaker anonymization suppresses speaker-identifying attributes from speech while preserving linguistic content and quality. We propose repurposing XTTSv2, a multilingual voice cloning model trained on 27k hours of speech, for speaker anonymization without retraining. Our key insight is that XTTSv2's voice cloning capabilities preserve prosodic structure independently of speaker identity, enabling voice conversion by conditioning on a pseudo-speaker. We introduce an iterative refinement strategy that balances privacy and utility by maximizing a harmonic mean of speaker dissimilarity and intelligibility. Evaluated on seven European languages across CommonVoice and Multilingual LibriSpeech, our system achieves near-optimal privacy (EER $\approx$ 0.49), competitive intelligibility, and substantially better speech quality than dedicated anonymization baselines, while requiring no language-specific training. We release the code here: https://github.com/rm00cr/coqui-tts.
Explore related subjects
Keep this discovery
Romolo Muletta, Felix Matthias Saaro, Mark Cieliebak, Jan Deriu. 2026-08-27. Your Voice Cloning System is Secretly a Voice Anonymizer. https://arxiv.org/abs/2608.27360
Cite the original work for its findings. Save a collection to share your selection of sources.