SearcharxivSearch

arXiv subjects

Hui-Yin Wu

Publications and source records attributed to Hui-Yin Wu.

7 recordsLinked to original sources

MObyGaze: a film dataset of multimodal objectification densely annotated by experts

Characterizing and quantifying gender representation disparities in audiovisual storytelling contents is necessary to grasp how stereotypes may perpetuate on screen. In this article, we consider the high-level construct of objectification and introduce a new AI task to the ML community: characterize and quantify complex multimodal (visual, speech, audio) temporal patterns producing objectification in films. Building on film studies and psychology, we define the construct of objectification in a structured thesaurus involving 5 sub-constructs manifesting through 11 concepts spanning 3 modalities. We introduce the Multimodal Objectifying Gaze (MObyGaze) dataset, made of 20 movies annotated densely by experts for objectification levels and concepts over freely delimited segments: it amounts to 6072 segments over 43 hours of video with fine-grained localization and categorization. We formulate new video interpretation tasks, show the feasibility of multimodal objectification detection, and analyze data and model bias to propose improvements. We exemplify two applications of MObyGaze, showing how rich concept annotation can improve model reliability and explainability. We make our code and our dataset available to the community and described in the Croissant format: https://github.com/husky-helen/MObyGaze.

cs.CV

Quantifying the Cost of Manual Navigation: A Comparison of Gesture-Based Magnification versus Direct Access Reading in Digital Layout-based Documents

Understanding how diverse audiences engage with structured media is critical to ensure a consistent quality of experience. In this context, we quantify the behavioral and performance cost of manual navigation (e.g., pinch and zoom) versus direct structural access in layout-based digital documents. We specifically investigate newspaper reading when visual access to structural cues (headlines as entry points) is constrained. Participants completed two tasks-reading all headlines aloud and locating target articles-under two conditions: (1) original edition with gesture-based magnification (pan and zoom), which is the industry standard for digital documents, and (2) large-print edition supporting direct-access reading. We collected performance measures (success ratio and completion time), behavioral integrity through reading path analysis, alongside perceived workload and preferences (NASA-TLX). Results from linear mixed-effects models show that the large-print condition yielded not only better performance than gesture-based magnification (18% improvement in reading speed, 30% improvement in speed to locate a target), but more importantly, restored the natural reading strategy that gesture-based magnification interaction disrupts. Readers also reported lower workload and higher preference. These findings highlight the importance of developing automated methods for generating large-print editions, where layout adaptation complements font scaling to support accessibility and quality of experience.

cs.HC

Leveraging multimodal explanatory annotations for video interpretation with Modality Specific Dataset

We examine the impact of concept-informed supervision on multimodal video interpretation models using MOByGaze, a dataset containing human-annotated explanatory concepts. We introduce Concept Modality Specific Datasets (CMSDs), which consist of data subsets categorized by the modality (visual, textual, or audio) of annotated concepts. Models trained on CMSDs outperform those using traditional legacy training in both early and late fusion approaches. Notably, this approach enables late fusion models to achieve performance close to that of early fusion models. These findings underscore the importance of modality-specific annotations in developing robust, self-explainable video models and contribute to advancing interpretable multimodal learning in complex video analysis.

cs.CV

DiVR: incorporating context from diverse VR scenes for human trajectory prediction

Virtual environments provide a rich and controlled setting for collecting detailed data on human behavior, offering unique opportunities for predicting human trajectories in dynamic scenes. However, most existing approaches have overlooked the potential of these environments, focusing instead on static contexts without considering userspecific factors. Employing the CREATTIVE3D dataset, our work models trajectories recorded in virtual reality (VR) scenes for diverse situations including road-crossing tasks with user interactions and simulated visual impairments. We propose Diverse Context VR Human Motion Prediction (DiVR), a cross-modal transformer based on the Perceiver architecture that integrates both static and dynamic scene context using a heterogeneous graph convolution network. We conduct extensive experiments comparing DiVR against existing architectures including MLP, LSTM, and transformers with gaze and point cloud context. Additionally, we also stress test our model's generalizability across different users, tasks, and scenes. Results show that DiVR achieves higher accuracy and adaptability compared to other models and to static graphs. This work highlights the advantages of using VR datasets for context-aware human trajectory modeling, with potential applications in enhancing user experiences in the metaverse. Our source code is publicly available at https://gitlab.inria.fr/ffrancog/creattive3d-divr-model.

cs.AI

Visual Objectification in Films: Towards a New AI Task for Video Interpretation

In film gender studies, the concept of 'male gaze' refers to the way the characters are portrayed on-screen as objects of desire rather than subjects. In this article, we introduce a novel video-interpretation task, to detect character objectification in films. The purpose is to reveal and quantify the usage of complex temporal patterns operated in cinema to produce the cognitive perception of objectification. We introduce the ObyGaze12 dataset, made of 1914 movie clips densely annotated by experts for objectification concepts identified in film studies and psychology. We evaluate recent vision models, show the feasibility of the task and where the challenges remain with concept bottleneck models. Our new dataset and code are made available to the community.

cs.CV

On-Line Cluster Reconstruction Of GEM Detector Based On FPGA Technology

In this work, a serial on-line cluster reconstruction technique based on FPGA technology was developed to compress experiment data and reduce the dead time of data transmission and storage. At the same time, X-ray imaging experiment based on a two-dimensional positive sensitive triple GEM detector with an effective readout area of 10 cm*10 cm was done to demonstrate this technique with FPGA development board. The result showed that the reconstruction technology was practicality and efficient. It provides a new idea for data compression of large spectrometers.

physics.ins-det

The Study of Cosmic Ray Tomography Using Multiple Scattering of Muons for Imaging of High-Z Materials

Muon tomography is developing as a promising system to detect high-Z (atomic number) material for ensuring homeland security. In the present work, three kinds of spatial locations of materials which are made of aluminum, iron, lead and uranium are simulated with GEANT4 codes, which are horizontal, diagonal and vertical objects, respectively. Two statistical algorithms are used with MATLAB software to reconstruct the image of detected objects, which are the Point of Closet Approach (PoCA) and Maximum Likelihood Scattering-Expectation Maximization iterative algorithm (MLS-EM), respectively. Two analysis methods are used to evaluate the quality of reconstruction image, which are the Receiver Operating Characteristic (ROC) and the localization ROC (LROC) curves, respectively. The reconstructed results show that, compared with PoCA algorithm, MLS-EM can achieve a better image quality in both edge preserving and noise reduction. And according to the analysis of ROC (LROC) curves, it shows that MLS-EM algorithm can discriminate and exclude the presence and location of high-Z object with a high efficiency, which is more flexible with an different EM algorithm employed than prior work. Furthermore the MLS-EM iterative algorithm will be modified and ran in parallel executive way for improving the reconstruction speed.

physics.ins-det