arXiv · 2512.05121
PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional Styles
Abstract
PESTalk is a novel method for generating 3D facial animations with personalized emotional styles directly from speech. It overcomes key limitations of existing approaches by introducing a Dual-Stream Emotion Extractor (DSEE) that captures both time and frequency-domain audio features for fine-grained emotion analysis, and an Emotional Style Modeling Module (ESMM) that models individual expression patterns based on voiceprint characteristics. To address data scarcity, the method leverages a newly constructed 3D-EmoStyle dataset. Evaluations demonstrate that PESTalk outperforms state-of-the-art methods in producing realistic and personalized facial animations.
Explore related subjects
Keep this discovery
Tianshun Han, Benjia Zhou, Ajian Liu, Yanyan Liang, Du Zhang, Zhen Lei, Jun Wan. 2025-10-13. PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional Styles. https://arxiv.org/abs/2512.05121
Cite the original work for its findings. Save a collection to share your selection of sources.