arXiv · 2609.29556
Cross-Modal Emotion Understanding: A Transformer-GAT Approach for Dialogue Emotion Recognition
Abstract
Multimodal emotion recognition is a key research area in affective computing, with applications in sentiment analysis, intelligent customer service, and human-computer interaction. However, existing methods often rely on single-modal features or simple multimodal fusion, failing to capture the synergy between global and local contexts, which limits model performance and emotion understanding. To address this challenge, we propose Transformer-GAT, a hybrid framework that combines Transformer and the Graph Attention Network to enable cross-modal emotion understanding. The Transformer is used to capture global semantic information, while the Graph Attention Network is employed to model fine-grained relationships between modalities, thereby enhancing the representation of emotional features. Experiments on the IEMOCAP and MELD datasets show that our model achieves weighted F1 scores of 72.45% and 77.37%, outperforming state-of-the-art methods. These results demonstrate that Transformer-GAT effectively integrates multimodal features, balances global and local contexts, and provides deeper emotional insights, offering new directions for multimodal emotion computing.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jiaqi Qiao, Yifan Lyu, Xiujuan Xu. 2026-09-02. Cross-Modal Emotion Understanding: A Transformer-GAT Approach for Dialogue Emotion Recognition. https://arxiv.org/abs/2609.29556
Cite the original work for its findings. Save a collection to share your selection of sources.