SearcharxivSearch

arXiv subjects

Eric J Gonzalez

Publications and source records attributed to Eric J Gonzalez.

8 recordsLinked to original sources

Generative Proxy: Synthesizing Proxy-Based Interfaces for Real-World Interaction Across AR Glasses

Interacting with real-world objects in AR is difficult, especially when targets are distant, cluttered, or occluded. These challenges are amplified on emerging lightweight AR glasses, which often lack binocular or large field of view on display, but also continuous inputs, such as hand or eye tracking. Proxy-based interfaces offer an alternative by allowing users to interact with virtual abstractions of physical objects that can be repositioned, reorganized, and adapted to the task and device. However, designing such interfaces is currently manual and highly device-specific. We present Generative Proxy, a method for automatically generating proxy-based interfaces from three specifications: scene, intent, and device capabilities. We formulate generation as a constrained synthesis problem that first produces valid interfaces for the target device and task, then ranks candidates using semantic and articulatory distance inspired by direct manipulation theory. We demonstrate Generative Proxy across diverse scenes, device profiles, and user intents. Expert evaluation shows initial evidence that generated proxy UIs are useful and usable, highlighting proxy-based abstraction as a promising interaction paradigm for future AR glasses.

cs.HC

VisionClaw: Always-On AI Agents through Smart Glasses

We present VisionClaw, an always-on wearable AI agent that integrates live egocentric perception with agentic task execution. Running on Meta Ray-Ban smart glasses, VisionClaw continuously perceives real-world context and enables in-situ, speech-driven action initiation and delegation via OpenClaw AI agents. Therefore, users can directly execute tasks through the smart glasses, such as adding real-world objects to an Amazon cart, generating notes from physical documents, receiving meeting briefings on the go, creating events from posters, or controlling IoT devices. We evaluate VisionClaw through a controlled laboratory study (N=12) and a longitudinal deployment study (N=5). Results show that integrating perception and execution enables faster task completion and reduces interaction overhead compared to non-always-on and non-agent baselines. Beyond performance gains, deployment findings reveal a shift in interaction: tasks are initiated opportunistically during ongoing activities, and execution is increasingly delegated rather than manually controlled. These results suggest a new paradigm for wearable AI agents, where perception and action are continuously coupled to support situated, hands-free interaction.

cs.HC

Navig-AI-tion: Navigation by Contextual AI and Spatial Audio

Audio-only walking navigation can leave users disoriented, relying on vague cardinal directions and lacking real-time environmental context, leading to frequent errors. To address this, we present a novel system that integrates a Vision Language Model (VLM) with a spatial audio cue. Our system extracts environmental landmarks to anchor navigation instructions and, crucially, provides a directional spatial audio signal when the user faces the wrong direction, indicating the precise turn direction. In a user study (n=12), the spatial audio cue with VLM reduced route deviations compared to both VLM-only and Google Maps (audio-only) baseline systems. Users reported that the spatial audio cue effectively supported orientation and that landmark-anchored instructions provided a better navigation experience over audio-only Google Maps. This work serves as an initial look at the utility of future audio-only navigation systems for incorporating directional cues, especially real-time corrective spatial audio.

cs.HC

Semantic Reality: Interactive Context-Aware Visualization of Inter-Object Relationships in Augmented Reality

Bridging the physical and digital world through interaction remains a core challenge in augmented reality (AR). Existing systems target single objects, limiting support for planning, comparison, and assembly tasks that depend on relationships among multiple items. We present Semantic Reality, an AR system focused on surfacing inter-object connectivity and making it interactive. Leveraging multimodal reasoning, spatial anchoring, and physical action recognition, Semantic Reality maintains a persistent model of objects around the user and their relationships. Connections are visualized in-situ to highlight compatibility, reveal next steps, and reduce ambiguity during tasks. We contribute a connectivity-centered interaction paradigm and a system architecture that couples anchor tracking, action sensing, and model inference to construct a live connectivity graph. In an exploratory study comparing Semantic Reality to a single-object baseline, participants reported clearer inter-object understanding and higher engagement and satisfaction, without increased workload. A scenario study illustrates where connectivity aids planning, sequencing, and disambiguation.

cs.HC

Break the Window: Exploring Spatial Decomposition of Webpages in XR

Most XR web browsers still present webpages as a single floating window, carrying over desktop design assumptions into immersive space. We explore an alternative by breaking the browser window and distributing a webpage into spatial UI chunks within a mixed-reality workspace. We present Break-the-Window (BTW), an exploratory prototype that spatially decomposes live, fully functional webpages into movable panels supporting mid-air and surface-attached placement, as well as direct touch and ray-based interaction. Through a formative study with XR practitioners and an exploratory qualitative study with 15 participants, we observed how spatial decomposition supports distributed attention and spatial meaning-making, while also surfacing challenges around coordination effort, interaction precision, and the lack of shared spatial UI conventions. This work invites discussion on how web interfaces might be reimagined for spatial computing beyond the single-window paradigm.

cs.HC

ForcePinch: Force-Responsive Spatial Interaction for Tracking Speed Control in XR

Spatial interaction in 3D environments requires balancing efficiency and precision, which requires dynamic tracking speed adjustments. However, existing techniques often couple tracking speed adjustments directly with hand movements, reducing interaction flexibility. Inspired by the natural friction control inherent in the physical world, we introduce ForcePinch, a novel force-responsive spatial interaction method that enables users to intuitively modulate pointer tracking speed and smoothly transition between rapid and precise movements by varying their pinching force. To implement this concept, we developed a hardware prototype integrating a pressure sensor with a customizable mapping function that translates pinching force into tracking speed adjustments. We conducted a user study with 20 participants performing well-established 1D, 2D, and 3D object manipulation tasks, comparing ForcePinch against the distance-responsive technique Go-Go and speed-responsive technique PRISM. Results highlight distinctive characteristics of the force-responsive approach across different interaction contexts. Drawing on these findings, we highlight the contextual meaning and versatility of force-responsive interactions through four illustrative examples, aiming to inform and inspire future spatial interaction design.

cs.HC

EmBARDiment: an Embodied AI Agent for Productivity in XR

XR devices running chat-bots powered by Large Language Models (LLMs) have the to become always-on agents that enable much better productivity scenarios. Current screen based chat-bots do not take advantage of the the full-suite of natural inputs available in XR, including inward facing sensor data, instead they over-rely on explicit voice or text prompts, sometimes paired with multi-modal data dropped as part of the query. We propose a solution that leverages an attention framework that derives context implicitly from user actions, eye-gaze, and contextual memory within the XR environment. Our work minimizes the need for engineered explicit prompts, fostering grounded and intuitive interactions that glean user insights for the chat-bot.

cs.HC

Hovering Over the Key to Text Input in XR

Virtual, Mixed, and Augmented Reality (XR) technologies hold immense potential for transforming productivity beyond PC. Therefore there is a critical need for improved text input solutions for XR. However, achieving efficient text input in these environments remains a significant challenge. This paper examines the current landscape of XR text input techniques, focusing on the importance of keyboards (both physical and virtual) as essential tools. We discuss the unique challenges and opportunities presented by XR, synthesizing key trends from existing solutions.

cs.HC