arXiv · 2508.07086
SEF-MK: Speaker-Embedding-Free Voice Anonymization through Multi-k-means Quantization
Abstract
Voice anonymization protects speaker privacy by concealing identity while preserving linguistic and paralinguistic content. Self-supervised learning (SSL) representations encode linguistic features but preserve speaker traits. We propose a novel speaker-embedding-free framework called SEF-MK. Instead of using a single k-means model trained on the entire dataset, SEF-MK anonymizes SSL representations for each utterance by randomly selecting one of multiple k-means models, each trained on a different subset of speakers. We explore this approach from both attacker and user perspectives. Extensive experiments show that, compared to a single k-means model, SEF-MK with multiple k-means models better preserves linguistic and emotional content from the user's viewpoint. However, from the attacker's perspective, utilizing multiple k-means models boosts the effectiveness of privacy attacks. These insights can aid users in designing voice anonymization systems to mitigate attacker threats.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Beilong Tang, Xiaoxiao Miao, Xin Wang, Ming Li. 2025-08-09. SEF-MK: Speaker-Embedding-Free Voice Anonymization through Multi-k-means Quantization. https://arxiv.org/abs/2508.07086
Cite the original work for its findings. Save a collection to share your selection of sources.