SearcharxivSearch

arXiv subjects

Vitaliy Popov

Publications and source records attributed to Vitaliy Popov.

8 recordsLinked to original sources

COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention

COSI-Lab presents a multimodal, multi-sensor dataset of an interdisciplinary scientific workshop containing 32 academics at an international conference. It captures ecologically valid social interactions in a weakly scripted setting consisting of two 30-minute mingling sessions with real professional and social consequences for the participants involved. We argue that future intelligent systems could be better equipped to handle subjective perceptions by modeling their multiplicity not as label noise but as a explainable perspective-driven reasoning process. We focus on the Apparent Intent Inference (AII) problem as determined by ex-situ observers and conceptualize intentions to be independent of manifest future outcomes. We contribute 1. a novel annotation process for AII that accounts for a perceiver's own interpretative tendencies, 2. quantitative and qualitative analyses of intent narratives with respect to diversity, grounding, and plausibility; 3. benchmark tasks for AII and surrounding relevant contextual factors such as social involvement; 4. speech quality audio for all participants as well as privacy preserving multi-modal data, enabling lexical and nonverbal behavior analysis; and 5. coupling of self-reported goals of each participant (30 minute to 3 hour) with annotated AII (seconds).

cs.HC

Towards Modeling Situational Awareness Through Visual Attention in Clinical Simulations

Situational awareness (SA) is essential for effective team performance in time-critical clinical environments, yet its dynamic and distributed nature remains difficult to characterize. In this preliminary study, we apply Transition Network Analysis (TNA) to model visual attention in multiperson VR-based cardiac arrest simulations. Using eye-tracking data from 40 clinicians assigned to four standardized roles (Airway, CPR, Defib, TeamLead), we construct gaze transition networks between clinically meaningful areas of interest (AOIs) and extract metrics such as entropy and self-loop rate to quantify attentional structure and flow. Our findings reveal that individual and team's visual attention is dynamically and adaptively redistributed across roles and scenario phases, with those in CPR roles narrowing their focus to execution-critical tasks and those in the TeamLead role concentrating on global monitoring as clinical demands evolve. TNA thus provides a powerful lens for mapping functional differentiation of team cognition and may support the development of phase-sensitive analytics and targeted instructional interventions in acute care training.

cs.HC

CliniDial: A Naturally Occurring Multimodal Dialogue Dataset for Team Reflection in Action During Clinical Operation

In clinical operations, teamwork can be the crucial factor that determines the final outcome. Prior studies have shown that sufficient collaboration is the key factor that determines the outcome of an operation. To understand how the team practices teamwork during the operation, we collected CliniDial from simulations of medical operations. CliniDial includes the audio data and its transcriptions, the simulated physiology signals of the patient manikins, and how the team operates from two camera angles. We annotate behavior codes following an existing framework to understand the teamwork process for CliniDial. We pinpoint three main characteristics of our dataset, including its label imbalances, rich and natural interactions, and multiple modalities, and conduct experiments to test existing LLMs' capabilities on handling data with these characteristics. Experimental results show that CliniDial poses significant challenges to the existing models, inviting future effort on developing methods that can deal with real-world clinical data. We open-source the codebase at https://github.com/MichiganNLP/CliniDial

cs.CL

Non-linear, Team-based VR Training for Cardiac Arrest Care with enhanced CRM Toolkit

This paper introduces iREACT, a novel VR simulation addressing key limitations in traditional cardiac arrest (CA) training. Conventional methods struggle to replicate the dynamic nature of real CA events, hindering Crew Resource Management (CRM) skill development. iREACT provides a non-linear, collaborative environment where teams respond to changing patient states, mirroring real CA complexities. By capturing multi-modal data (user actions, cognitive load, visual gaze) and offering real-time and post-session feedback, iREACT enhances CRM assessment beyond traditional methods. A formative evaluation with medical experts underscores its usability and educational value, with potential applications in other high-stakes training scenarios to improve teamwork, communication, and decision-making.

cs.HC

eXplainMR: Generating Real-time Textual and Visual eXplanations to Facilitate UltraSonography Learning in MR

eXplainMR is a Mixed Reality tutoring system designed for basic cardiac surface ultrasound training. Trainees wear a head-mounted display (HMD) and hold a controller, mimicking a real ultrasound probe, while treating a desk surface as the patient's body for low-cost and anywhere training. eXplainMR engages trainees with troubleshooting questions and provides automated feedback through four key mechanisms: 1) subgoals that break down tasks into single-movement steps, 2) textual explanations comparing the current incorrect view with the target view, 3) real-time segmentation and annotation of ultrasound images for direct visualization, and 4) the 3D visual cues provide further explanations on the intersection between the slicing plane and anatomies.

cs.HC

Looking Together $\neq$ Seeing the Same Thing: Understanding Surgeons' Visual Needs During Intra-operative Coordination and Instruction

Shared gaze visualizations have been found to enhance collaboration and communication outcomes in diverse HCI scenarios including computer supported collaborative work and learning contexts. Given the importance of gaze in surgery operations, especially when a surgeon trainer and trainee need to coordinate their actions, research on the use of gaze to facilitate intra-operative coordination and instruction has been limited and shows mixed implications. We performed a field observation of 8 surgeries and an interview study with 14 surgeons to understand their visual needs during operations, informing ways to leverage and augment gaze to enhance intra-operative coordination and instruction. We found that trainees have varying needs in receiving visual guidance which are often unfulfilled by the trainers' instructions. It is critical for surgeons to control the timing of the gaze-based visualizations and effectively interpret gaze data. We suggest overlay technologies, e.g., gaze-based summaries and depth sensing, to augment raw gaze in support of surgical coordination and instruction.

cs.HC

Surgment: Segmentation-enabled Semantic Search and Creation of Visual Question and Feedback to Support Video-Based Surgery Learning

Videos are prominent learning materials to prepare surgical trainees before they enter the operating room (OR). In this work, we explore techniques to enrich the video-based surgery learning experience. We propose Surgment, a system that helps expert surgeons create exercises with feedback based on surgery recordings. Surgment is powered by a few-shot-learning-based pipeline (SegGPT+SAM) to segment surgery scenes, achieving an accuracy of 92\%. The segmentation pipeline enables functionalities to create visual questions and feedback desired by surgeons from a formative study. Surgment enables surgeons to 1) retrieve frames of interest through sketches, and 2) design exercises that target specific anatomical components and offer visual feedback. In an evaluation study with 11 surgeons, participants applauded the search-by-sketch approach for identifying frames of interest and found the resulting image-based questions and feedback to be of high educational value.

cs.HC

Measurement of $D^0-\bar{D}^0$ mixing parameters using semileptonic decays of neutral kaon

We propose a new method to extract $D^0-\bar{D}^0$ mixing parameters using the $D^0 \to \bar{K}{}^0 π^0$ decay with the $\bar{K}{}^0$ reconstructed in the semileptonic mode. Although a $K^0 \to π^\pm l^\mp ν_l$ decay suffers from low statistics and complexity of the secondary vertex reconstruction in comparison to the standard $K_{S} \to π^+ π^-$ vertex, it provides much richer, sometimes unique information about the initial state of a $K^0$-meson produced in a heavy-flavor hadron decay. In this paper it is shown that the reconstruction of the chain $D^0 \to K^0 (π^\pm l^\mp ν_l) π^0$ allows one to extract the strong phase difference between the doubly Cabibbo-suppressed and Cabibbo-favored decay amplitudes, which is of key importance for determination of the $D^0-\bar{D}^0$ mixing parameters.

hep-ph