arXiv · 2302.06227
Fast and small footprint Hybrid HMM-HiFiGAN based system for speech synthesis in Indian languages
Abstract
Hidden-Markov-model (HMM) based text-to-speech (HTS) offers flexibility in speaking styles along with fast training and synthesis while being computationally less intense. HTS performs well even in low-resource scenarios. The primary drawback is that the voice quality is poor compared to that of E2E systems. A hybrid approach combining HMM-based feature generation and neural-network-based HiFi-GAN vocoder to improve HTS synthesis quality is proposed. HTS is trained on high-resolution mel-spectrograms instead of conventional mel generalized coefficients (MGC), and the output mel-spectrogram corresponding to the input text is used in a HiFi-GAN vocoder trained on Indic languages, to produce naturalness that is equivalent to that of E2E systems, as evidenced from the DMOS and PC tests.
Explore related subjects
Keep this discovery
Sudhanshu Srivastava, Ishika Gupta, Anusha Prakash, Jom Kuriakose, Hema A. Murthy. 2023-02-13. Fast and small footprint Hybrid HMM-HiFiGAN based system for speech synthesis in Indian languages. https://arxiv.org/abs/2302.06227
Cite the original work for its findings. Save a collection to share your selection of sources.