arXiv · 2211.08978
Rapid Connectionist Speaker Adaptation
Abstract
We present SVCnet, a system for modelling speaker variability. Encoder Neural Networks specialized for each speech sound produce low dimensionality models of acoustical variation, and these models are further combined into an overall model of voice variability. A training procedure is described which minimizes the dependence of this model on which sounds have been uttered. Using the trained model (SVCnet) and a brief, unconstrained sample of a new speaker's voice, the system produces a Speaker Voice Code that can be used to adapt a recognition system to the new speaker without retraining. A system which combines SVCnet with an MS-TDNN recognizer is described
Explore related subjects
Keep this discovery
Michael Witbrock, Patrick Haffner. 2022-11-15. Rapid Connectionist Speaker Adaptation. https://doi.org/10.1109/icassp.1992.225874
Cite the original work for its findings. Save a collection to share your selection of sources.