arXiv · 1506.05268
Deep Denoising Auto-encoder for Statistical Speech Synthesis
Abstract
This paper proposes a deep denoising auto-encoder technique to extract better acoustic features for speech synthesis. The technique allows us to automatically extract low-dimensional features from high dimensional spectral features in a non-linear, data-driven, unsupervised way. We compared the new stochastic feature extractor with conventional mel-cepstral analysis in analysis-by-synthesis and text-to-speech experiments. Our results confirm that the proposed method increases the quality of synthetic speech in both experiments.
Explore related subjects
Keep this discovery
Zhenzhou Wu, Shinji Takaki, Junichi Yamagishi. 2015-06-17. Deep Denoising Auto-encoder for Statistical Speech Synthesis. https://arxiv.org/abs/1506.05268
Cite the original work for its findings. Save a collection to share your selection of sources.