SearcharxivSearch

arXiv subjects

Christina Sartzetaki

Publications and source records attributed to Christina Sartzetaki.

2 recordsLinked to original sources

A large dataset of human EEG responses to short naturalistic videos for studying dynamic visual event processing

Vision neuroscience has experienced a surge in the collection and use of large-scale datasets of brain responses to naturalistic images. However, static images lack the temporal dimension essential for understanding how vision is solved in the brain during dynamic real life settings. To facilitate the study of the neural correlates of dynamic visual event perception, we introduce the EEG Moments Dataset (EMD). EMD consists of 128-channel EEG responses and eye-tracking recordings of 6 human participants viewing 1,102 short naturalistic videos (3-second long; with audio track) while maintaining central fixation. We show that EMD's EEG responses well encode stimulus-related information, exhibit a temporal correspondence with the video stimuli, and have a rich representational content revealed by brain encoding models based on different feature spaces. Furthermore, complemented by the BOLD Moments Dataset (BMD) - an existing large-scale dataset of human functional magnetic resonance imaging (fMRI) responses for the same videos - EMD enables spatio-temporally resolved investigations of brain responses to dynamic visual events. We release EMD's EEG and eye-tracking data in both raw and preprocessed format, along with the 1,102 video stimuli, and rich stimulus metadata. Finally, we provide an interactive code tutorial to familiarize with EMD's preprocessed data, stimuli, and stimulus metadata.

q-bio.NC

Extending Compositional Attention Networks for Social Reasoning in Videos

We propose a novel deep architecture for the task of reasoning about social interactions in videos. We leverage the multi-step reasoning capabilities of Compositional Attention Networks (MAC), and propose a multimodal extension (MAC-X). MAC-X is based on a recurrent cell that performs iterative mid-level fusion of input modalities (visual, auditory, text) over multiple reasoning steps, by use of a temporal attention mechanism. We then combine MAC-X with LSTMs for temporal input processing in an end-to-end architecture. Our ablation studies show that the proposed MAC-X architecture can effectively leverage multimodal input cues using mid-level fusion mechanisms. We apply MAC-X to the task of Social Video Question Answering in the Social IQ dataset and obtain a 2.5% absolute improvement in terms of binary accuracy over the current state-of-the-art.

cs.CV