arXiv · 2609.30517
Seeing Speech: Learning Visible Articulatory Dynamics for Speech-Driven 3D Facial Animation
Abstract
Recent progress in speech-driven 3D facial animation has improved vertex-level reconstruction quality, but speech-consistent visible articulation remains difficult. This is because speech production follows structured and constrained articulators' coordination and the mapping from acoustics to motion is inherently one-to-many. Motivated by the structured patterns of visible articulation, we propose a novel articulation-aware framework that models visible speech through directional articulatory motions and composes them into surface-consistent 3D facial motion. To represent visible articulation with three directional articulatory motions, spreading, opening, and protrusion, we propose a Speech--Articulatory Memory (SAM) that captures the correspondence between speech and these motions under phonetic context through retrieval and decoding based on a key-value memory structure. Then, a Topology-aware Articulatory Composition (TAC) integrates the predicted directional articulatory motions under mesh topology to produce surface-consistent 3D facial motion. Experiments on VOCASET and TFHP show that our method achieves state-of-the-art performance on standard reconstruction metrics and improves visible articulatory distance and velocity errors for lip articulation, while a user study confirms clear preference in lip sync and realism.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hyung Kyu Kim, Byungchan Hwang, Hak Gu Kim. 2026-09-24. Seeing Speech: Learning Visible Articulatory Dynamics for Speech-Driven 3D Facial Animation. https://arxiv.org/abs/2609.30517
Cite the original work for its findings. Save a collection to share your selection of sources.