arXiv · 2505.13688
Gaze-Enhanced Multimodal Turn-Taking Prediction in Triadic Conversations
Abstract
Turn-taking prediction is crucial for seamless interactions. This study introduces a novel, lightweight framework for accurate turn-taking prediction in triadic conversations without relying on computationally intensive methods. Unlike prior approaches that either disregard gaze or treat it as a passive signal, our model integrates gaze with speaker localization, structuring it within a spatial constraint to transform it into a reliable predictive cue. Leveraging egocentric behavioral cues, our experiments demonstrate that incorporating gaze data from a single-user significantly improves prediction performance, while gaze data from multiple-users further enhances it by capturing richer conversational dynamics. This study presents a lightweight and privacy-conscious approach to support adaptive, directional sound control, enhancing speech intelligibility in noisy environments, particularly for hearing assistance in smart glasses.
Explore related subjects
Keep this discovery
Seongsil Heo, Calvin Murdock, Michael Proulx, Christi Miller. 2025-05-19. Gaze-Enhanced Multimodal Turn-Taking Prediction in Triadic Conversations. https://arxiv.org/abs/2505.13688
Cite the original work for its findings. Save a collection to share your selection of sources.