arXiv · 1906.05682
Focal Loss based Residual Convolutional Neural Network for Speech Emotion Recognition
Abstract
This paper proposes a Residual Convolutional Neural Network (ResNet) based on speech features and trained under Focal Loss to recognize emotion in speech. Speech features such as Spectrogram and Mel-frequency Cepstral Coefficients (MFCCs) have shown the ability to characterize emotion better than just plain text. Further Focal Loss, first used in One-Stage Object Detectors, has shown the ability to focus the training process more towards hard-examples and down-weight the loss assigned to well-classified examples, thus preventing the model from being overwhelmed by easily classifiable examples.
Explore related subjects
Keep this discovery
Suraj Tripathi, Abhay Kumar, Abhiram Ramesh, Chirag Singh, Promod Yenigalla. 2019-06-11. Focal Loss based Residual Convolutional Neural Network for Speech Emotion Recognition. https://arxiv.org/abs/1906.05682
Cite the original work for its findings. Save a collection to share your selection of sources.