arXiv · 2409.08913
HLTCOE JHU Submission to the Voice Privacy Challenge 2024
Abstract
We present a number of systems for the Voice Privacy Challenge, including voice conversion based systems such as the kNN-VC method and the WavLM voice Conversion method, and text-to-speech (TTS) based systems including Whisper-VITS. We found that while voice conversion systems better preserve emotional content, they struggle to conceal speaker identity in semi-white-box attack scenarios; conversely, TTS methods perform better at anonymization and worse at emotion preservation. Finally, we propose a random admixture system which seeks to balance out the strengths and weaknesses of the two category of systems, achieving a strong EER of over 40% while maintaining UAR at a respectable 47%.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Henry Li Xinyuan, Zexin Cai, Ashi Garg, Kevin Duh, Leibny Paola García-Perera, Sanjeev Khudanpur, Nicholas Andrews, Matthew Wiesner. 2024-09-13. HLTCOE JHU Submission to the Voice Privacy Challenge 2024. https://arxiv.org/abs/2409.08913
Cite the original work for its findings. Save a collection to share your selection of sources.