arXiv · 2609.36624
DualTrack: Synchronized speech-gesture generation via symmetric coupling of pretrained priors
Abstract
Joint speech-gesture synthesis must coordinate two modalities despite limited paired data. Existing approaches often lack bidirectional interaction, have limited language coverage, or simplify body and finger representations. We present DualTrack, which couples pretrained speech and motion priors on a shared 12.5 Hz timeline. Causal adapters exchange previous-packet information, while current-state fusion coordinates the streams before they separately complete sixteen-codebook packets. We evaluate 43 BEAT2 recordings in four languages, with speakers held out from joint training and validation. On the shared English/Spanish inputs, without speech or motion prefixes, DualTrack achieves lower word error rate and full-motion Fréchet Gesture Distance, higher beat consistency and speech naturalness than the evaluated GELINA baseline.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yuanzhuo Hu, Zehan Liu, Xiaoyi Qin, Ming Li. 2026-09-29. DualTrack: Synchronized speech-gesture generation via symmetric coupling of pretrained priors. https://arxiv.org/abs/2609.36624
Cite the original work for its findings. Save a collection to share your selection of sources.