arXiv · 2603.05977
Activation Steering for Accent-Neutralized Zero-Shot Text-To-Speech
Abstract
Zero-shot Text-to-Speech (TTS) models can generate speech that captures both the voice timbre and accent of a reference speaker. However, disentangling these attributes and individually controlling the output accent remains challenging. In this study, we introduce a post-hoc and training-free approach to neutralize accent while preserving the speaker's original timbre, utilizing inference-time activation steering. We first extract layer-specific "steering vectors" offline, which are derived from the activation differences within the TTS model between accented and native speech. During inference, the steering vectors guide the model to produce accent-neutralized, timbre-preserving speech. Empirical results demonstrate that the proposed steering vectors effectively mitigate the output accent and exhibit strong generalizability to unseen accented speakers, offering a practical solution for accent-free voice cloning. We further confirm that speaker timbre is largely preserved via speaker embedding cosine similarity and speaker verification acceptance rate computed in an estimated accent-orthogonal subspace, as well as subjective listening tests. Audio demo: https://accentsteer.github.io/
Explore related subjects
Keep this discovery
Mu Yang, John H. L. Hansen. 2026-03-06. Activation Steering for Accent-Neutralized Zero-Shot Text-To-Speech. https://arxiv.org/abs/2603.05977
Cite the original work for its findings. Save a collection to share your selection of sources.