arXiv · 1901.03601
Advanced Rich Transcription System for Estonian Speech
Abstract
This paper describes the current TT\"U speech transcription system for Estonian speech. The system is designed to handle semi-spontaneous speech, such as broadcast conversations, lecture recordings and interviews recorded in diverse acoustic conditions. The system is based on the Kaldi toolkit. Multi-condition training using background noise profiles extracted automatically from untranscribed data is used to improve the robustness of the system. Out-of-vocabulary words are recovered using a phoneme n-gram based decoding subgraph and a FST-based phoneme-to-grapheme model. The system achieves a word error rate of 8.1% on a test set of broadcast conversations. The system also performs punctuation recovery and speaker identification. Speaker identification models are trained using a recently proposed weakly supervised training method.
Explore related subjects
Keep this discovery
Tanel Alumäe, Ottokar Tilk, Asadullah. 2019-01-11. Advanced Rich Transcription System for Estonian Speech. https://doi.org/10.3233/978-1-61499-912-6-1
Cite the original work for its findings. Save a collection to share your selection of sources.