arXiv · 2312.13603
Style Modeling for Multi-Speaker Articulation-to-Speech
Abstract
In this paper, we propose a neural articulation-to-speech (ATS) framework that synthesizes high-quality speech from articulatory signal in a multi-speaker situation. Most conventional ATS approaches only focus on modeling contextual information of speech from a single speaker's articulatory features. To explicitly represent each speaker's speaking style as well as the contextual information, our proposed model estimates style embeddings, guided from the essential speech style attributes such as pitch and energy. We adopt convolutional layers and transformer-based attention layers for our model to fully utilize both local and global information of articulatory signals, measured by electromagnetic articulography (EMA). Our model significantly improves the quality of synthesized speech compared to the baseline in terms of objective and subjective measurements in the Haskins dataset.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Miseul Kim, Zhenyu Piao, Jihyun Lee, Hong-Goo Kang. 2023-12-21. Style Modeling for Multi-Speaker Articulation-to-Speech. https://arxiv.org/abs/2312.13603
Cite the original work for its findings. Save a collection to share your selection of sources.