arXiv · 2609.21683
Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation
Abstract
Conversational speech depends on dialogue context and the listener's immediately preceding behavior. We propose ReACT-TTS, a two-stage framework that uses a one-second pre-response listener facial sequence to plan the next utterance's emotion and prosody before speech realization. On a strict dyadic MELD protocol, Temporal conditioning yields higher mean macro-F1 and VAD concordance than Text-only across ten seeds, while accuracy remains essentially unchanged. Ablations show that temporal modeling performs best among the visual variants and that an explicit early-to-late difference is unnecessary; correct listener reactions also outperform cyclic mismatches on average. In a contextual-appropriateness study with 20 speech researchers, 76% of judgments prefer Temporal, 9% Text-only, and 15% report no preference. We further connect the predicted response style to a Grad-TTS backbone for end-to-end speech realization. Overall, the results support pre-response listener dynamics as complementary cues for conversational response planning. The source code is available at https://github.com/CYJ1/ReACT-TTS_public.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yunji Chu. 2026-09-18. Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation. https://arxiv.org/abs/2609.21683
Cite the original work for its findings. Save a collection to share your selection of sources.