arXiv · 2409.05089
Leveraging WaveNet for Dynamic Listening Head Modeling from Speech
Abstract
The creation of listener facial responses aims to simulate interactive communication feedback from a listener during a face-to-face conversation. Our goal is to generate believable videos of listeners' heads that respond authentically to a single speaker by a sequence-to-sequence model with an combination of WaveNet and Long short-term memory network. Our approach focuses on capturing the subtle nuances of listener feedback, ensuring the preservation of individual listener identity while expressing appropriate attitudes and viewpoints. Experiment results show that our method surpasses the baseline models on ViCo benchmark Dataset.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Minh-Duc Nguyen, Hyung-Jeong Yang, Seung-Won Kim, Ji-Eun Shin, Soo-Hyung Kim. 2024-09-08. Leveraging WaveNet for Dynamic Listening Head Modeling from Speech. https://arxiv.org/abs/2409.05089
Cite the original work for its findings. Save a collection to share your selection of sources.