arXiv · 1807.00868
Exploring End-to-End Techniques for Low-Resource Speech Recognition
Abstract
In this work we present simple grapheme-based system for low-resource speech recognition using Babel data for Turkish spontaneous speech (80 hours). We have investigated different neural network architectures performance, including fully-convolutional, recurrent and ResNet with GRU. Different features and normalization techniques are compared as well. We also proposed CTC-loss modification using segmentation during training, which leads to improvement while decoding with small beam size. Our best model achieved word error rate of 45.8%, which is the best reported result for end-to-end systems using in-domain data for this task, according to our knowledge.
Explore related subjects
Keep this discovery
Vladimir Bataev, Maxim Korenevsky, Ivan Medennikov, Alexander Zatvornitskiy. 2018-07-02. Exploring End-to-End Techniques for Low-Resource Speech Recognition. https://arxiv.org/abs/1807.00868
Cite the original work for its findings. Save a collection to share your selection of sources.