arXiv · 2508.12854
E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
Abstract
Multimodal Empathetic Response Generation (MERG) is crucial for building emotionally intelligent human-computer interactions. Although large language models (LLMs) have improved text-based ERG, challenges remain in handling multimodal emotional content and maintaining identity consistency. Thus, we propose E3RG, an Explicit Emotion-driven Empathetic Response Generation System based on multimodal LLMs which decomposes MERG task into three parts: multimodal empathy understanding, empathy memory retrieval, and multimodal response generation. By integrating advanced expressive speech and video generative models, E3RG delivers natural, emotionally rich, and identity-consistent responses without extra training. Experiments validate the superiority of our system on both zero-shot and few-shot settings, securing Top-1 position in the Avatar-based Multimodal Empathy Challenge on ACM MM 25. Our code is available at https://github.com/RH-Lin/E3RG.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ronghao Lin, Shuai Shen, Weipeng Hu, Qiaolin He, Aolin Xiong, Li Huang, Haifeng Hu, Yap-peng Tan. 2025-08-18. E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model. https://arxiv.org/abs/2508.12854
Cite the original work for its findings. Save a collection to share your selection of sources.