arXiv · 2509.12295
More Similar than Dissimilar: Modeling Annotators for Cross-Corpus Speech Emotion Recognition
Abstract
Speech emotion recognition systems often predict a consensus value generated from the ratings of multiple annotators. However, these models have limited ability to predict the annotation of any one person. Alternatively, models can learn to predict the annotations of all annotators. Adapting such models to new annotators is difficult as new annotators must individually provide sufficient labeled training data. We propose to leverage inter-annotator similarity by using a model pre-trained on a large annotator population to identify a similar, previously seen annotator. Given a new, previously unseen, annotator and limited enrollment data, we can make predictions for a similar annotator, enabling off-the-shelf annotation of unseen data in target datasets, providing a mechanism for extremely low-cost personalization. We demonstrate our approach significantly outperforms other off-the-shelf approaches, paving the way for lightweight emotion adaptation, practical for real-world deployment.
Explore related subjects
Keep this discovery
James Tavernor, Emily Mower Provost. 2025-09-15. More Similar than Dissimilar: Modeling Annotators for Cross-Corpus Speech Emotion Recognition. https://arxiv.org/abs/2509.12295
Cite the original work for its findings. Save a collection to share your selection of sources.