arXiv · 2609.22031
Partial Accent-Control Editing in Frozen Speech Representations for Accent Conversion
Abstract
Accent conversion is the task of modifying a speech recording so that it sounds closer to a target accent while preserving linguistic content and other speaker-related characteristics. Most accent conversion systems use trained, generative models. Although they can induce target-accented speech, the strength of accent modification is not controllable at inference time, making it difficult to analyse how the strength of accent conversion affects source preservation. We propose Partial Accent-Control Editing (PACE), an accent conversion framework based upon the editing of frozen WavLM representations without the training of an accent-conditioned generator. A constrained edit is first applied to source WavLM features, which are then fused with target-accent reference features retrieved from non-parallel accent examples. Fusion weights are used to control the trade-off between accent conversion strength and the degradation of other source attributes. Using a suite of five metrics, and with the cost of accent-entangled speaker similarity, we show that PACE is substantially superior to a pair of competitive baselines in terms of both accent conversion and source preservation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yangyang Qu, Michele Panariello, Massimiliano Todisco, Nicholas Evans. 2026-09-18. Partial Accent-Control Editing in Frozen Speech Representations for Accent Conversion. https://arxiv.org/abs/2609.22031
Cite the original work for its findings. Save a collection to share your selection of sources.