SearcharxivSearch

arXiv subjects

John Dudley

Publications and source records attributed to John Dudley.

8 recordsLinked to original sources

Extended Reality as a Mediation Layer for Situated Human Control in Human-Robot Teaming

Extended Reality (XR) is increasingly used in human-robot interaction to communicate robot intent, planned motion, reachability, and state. We argue that XR should also be understood as a mediation layer for situated human control in human-robot teaming. Situated human control denotes the human collaborator's ability to understand, shape, authorize, and interrupt robot action within the concrete physical, social, and temporal context in which that action unfolds. We ground this perspective in scenarios from robot-assisted bedside nursing, multi-arm supervisory control, and collaborative assembly under divided attention. Across these scenarios, robot autonomy must remain inspectable and adjustable as people move, goals change, sensing is incomplete, control roles shift, and plans become invalid. We identify four mediation functions connecting human intent and robot autonomy, robot plans and human judgment, levels of shared control, and team roles, handover, and recovery. Building on these functions, we derive six design dimensions: joint action possibilities, socio-physical constraints, uncertainty and plan validity, multimodal control and correction, roles, handover, and accountability, and anticipatory recovery. The paper outlines a research agenda for XR systems that make robot autonomy more actionable and accountable in dynamic shared environments.

cs.HC

ReachVox: Clutter-free Reachability Visualization for Robot Motion Planning in Virtual Reality

Human-Robot-Collaboration can enhance workflows by leveraging the mutual strengths of human operators and robots. Planning and understanding robot movements remain major challenges in this domain. This problem is prevalent in dynamic environments that might need constant robot motion path adaptation. In this paper, we investigate whether a minimalistic encoding of the reachability of a point near an object of interest, which we call ReachVox, can aid the collaboration between a remote operator and a robotic arm in VR. Through a user study (n=20), we indicate the strength of the visualization relative to a point-based reachability check-up.

cs.HC

EnVisionVR: A Scene Interpretation Tool for Visual Accessibility in Virtual Reality

Effective visual accessibility in Virtual Reality (VR) is crucial for Blind and Low Vision (BLV) users. However, designing visual accessibility systems is challenging due to the complexity of 3D VR environments and the need for techniques that can be easily retrofitted into existing applications. While prior work has studied how to enhance or translate visual information, the advancement of Vision Language Models (VLMs) provides an exciting opportunity to advance the scene interpretation capability of current systems. This paper presents EnVisionVR, an accessibility tool for VR scene interpretation. Through a formative study of usability barriers, we confirmed the lack of visual accessibility features as a key barrier for BLV users of VR content and applications. In response, we designed and developed EnVisionVR, a novel visual accessibility system leveraging a VLM, voice input and multimodal feedback for scene interpretation and virtual object interaction in VR. An evaluation with 12 BLV users demonstrated that EnVisionVR significantly improved their ability to locate virtual objects, effectively supporting scene understanding and object interaction.

cs.HC

The Imaginative Generative Adversarial Network: Automatic Data Augmentation for Dynamic Skeleton-Based Hand Gesture and Human Action Recognition

Deep learning approaches deliver state-of-the-art performance in recognition of spatiotemporal human motion data. However, one of the main challenges in these recognition tasks is limited available training data. Insufficient training data results in over-fitting and data augmentation is one approach to address this challenge. Existing data augmentation strategies based on scaling, shifting and interpolating offer limited generalizability and typically require detailed inspection of the dataset as well as hundreds of GPU hours for hyperparameter optimization. In this paper, we present a novel automatic data augmentation model, the Imaginative Generative Adversarial Network (GAN), that approximates the distribution of the input data and samples new data from this distribution. It is automatic in that it requires no data inspection and little hyperparameter tuning and therefore it is a low-cost and low-effort approach to generate synthetic data. We demonstrate our approach on small-scale skeleton-based datasets with a comprehensive experimental analysis. Our results show that the augmentation strategy is fast to train and can improve classification accuracy for both conventional neural networks and state-of-the-art methods.

cs.CV

Encode-Store-Retrieve: Augmenting Human Memory through Language-Encoded Egocentric Perception

We depend on our own memory to encode, store, and retrieve our experiences. However, memory lapses can occur. One promising avenue for achieving memory augmentation is through the use of augmented reality head-mounted displays to capture and preserve egocentric videos, a practice commonly referred to as lifelogging. However, a significant challenge arises from the sheer volume of video data generated through lifelogging, as the current technology lacks the capability to encode and store such large amounts of data efficiently. Further, retrieving specific information from extensive video archives requires substantial computational power, further complicating the task of quickly accessing desired content. To address these challenges, we propose a memory augmentation agent that involves leveraging natural language encoding for video data and storing them in a vector database. This approach harnesses the power of large vision language models to perform the language encoding process. Additionally, we propose using large language models to facilitate natural language querying. Our agent underwent extensive evaluation using the QA-Ego4D dataset and achieved state-of-the-art results with a BLEU score of 8.3, outperforming conventional machine learning models that scored between 3.4 and 5.8. Additionally, we conducted a user study in which participants interacted with the human memory augmentation agent through episodic memory and open-ended questions. The results of this study show that the agent results in significantly better recall performance on episodic memory tasks compared to human participants. The results also highlight the agent's practical applicability and user acceptance.

cs.CV

Akhmediev breather signatures from dispersive propagation of a periodically phase-modulated continuous wave

We investigate in detail the qualitative similarities between the pulse localization characteristics observed using sinusoidal phase modulation during linear propagation and those seen during the evolution of Akhmediev breathers during propagation in a system governed by the nonlinear Schr{ö}dinger equation. The profiles obtained at the point of maximum focusing indeed present very close temporal and spectral features. If the respective linear and nonlinear longitudinal evolutions of those profiles are similar in the vicinity of the point of maximum focusing, they may diverge significantly for longer propagation distance. Our analysis and numerical simulations are confirmed by experiments performed in optical fiber.

physics.optics

Phase evolution of Peregrine-like breathers in optics and hydrodynamics

We present a detailed study of the phase properties of rational breather waves observed in the hydrodynamic and optical domains, namely the Peregrine soliton and related second-order solution. At the point of maximum compression, our experimental results recorded in a wave tank or using an optical fiber platform reveal a characteristic phase shift that is multiple of $π$ between the central part of the pulse and the continuous background, in agreement with analytical and numerical predictions. We also stress the existence of a large longitudinal phase shift across the point of maximum compression.

physics.optics

Optical Rogue Waves in Whispering-Gallery-Mode Resonators

We report a theoretical study showing that rogue waves can emerge in whispering gallery mode resonators as the result of the chaotic interplay between Kerr nonlinearity and anomalous group-velocity dispersion. The nonlinear dynamics of the propagation of light in a whispering gallery-mode resonator is investigated using the Lugiato-Lefever equation, and we evidence a range of parameters where rare and extreme events associated with a non-gaussian statistics of the field maxima are observed.

physics.optics