arXiv · 2311.16484
Seeing Eye to AI: Comparing Human Gaze and Model Attention in Video Memorability
Abstract
Understanding what makes a video memorable has important applications in advertising or education technology. Towards this goal, we investigate spatio-temporal attention mechanisms underlying video memorability. Different from previous works that fuse multiple features, we adopt a simple CNN+Transformer architecture that enables analysis of spatio-temporal attention while matching state-of-the-art (SoTA) performance on video memorability prediction. We compare model attention against human gaze fixations collected through a small-scale eye-tracking study where humans perform the video memory task. We uncover the following insights: (i) Quantitative saliency metrics show that our model, trained only to predict a memorability score, exhibits similar spatial attention patterns to human gaze, especially for more memorable videos. (ii) The model assigns greater importance to initial frames in a video, mimicking human attention patterns. (iii) Panoptic segmentation reveals that both (model and humans) assign a greater share of attention to things and less attention to stuff as compared to their occurrence probability.
Explore related subjects
Keep this discovery
Prajneya Kumar, Eshika Khandelwal, Makarand Tapaswi, Vishnu Sreekumar. 2023-11-26. Seeing Eye to AI: Comparing Human Gaze and Model Attention in Video Memorability. https://arxiv.org/abs/2311.16484
Cite the original work for its findings. Save a collection to share your selection of sources.