SearcharxivSearch

arXiv subjects

Yoichi Ochiai

Publications and source records attributed to Yoichi Ochiai.

At least 19 recordsLinked to original sources

Event-Based Spatial-Carrier Interferometry for Surface-Normal Vibration-Waveform Reconstruction

Non-contact measurement of small vibrations perpendicular to a surface supports the evaluation of mechanical structures, but in camera-based interferometry, increasing the frame rate makes a trade-off with the field of view and spatial resolution. By recording only brightness changes, event cameras avoid this trade-off and reach high temporal and spatial resolution; our previously reported event topology-based visual vibrometer recovers vibration from apparent motion. This high-speed, high-resolution sensing is well suited to full-field measurement, yet such vibration produces too little apparent motion to capture its waveform. Here we show that event-based spatial-carrier interferometry reconstructs that waveform from moving interference fringes. That displacement moves the fringes, and signed event-density maps built from the event stream are demodulated at the spatial carrier to recover the interferometric phase and fix the otherwise ambiguous motion direction at turning points. Reconstructed waveforms agree with laser Doppler vibrometry over broad drive-frequency and amplitude ranges, with limits set by the maximum fringe speed and the sensor performance. Reconstruction is limited by a minimum aperture of about two fringe periods along the carrier and one along the fringes, which allows the surface to be mapped region by region. These results provide an empirical basis for full-field, spatially resolved interferometric vibrometry with event cameras as a non-contact measurement technique.

physics.app-ph

Galvanic Vestibular Stimulation in Latent Space

Galvanic vestibular stimulation (GVS) is widely used to modulate self-orientation, balance, and motion perception; the discriminability of frequency-encoded cues further suggests its potential as a standalone modality for embodied feedback. However, synthesizing GVS waveforms congruent with target events or bodily states remains challenging. GVS waveforms combine current direction, intensity, duration, and onset and offset transitions, yet how these parameters jointly shape users' perceptual and associative responses remains underexplored. To address this gap, we contribute a dataset linking GVS waveforms to free-form experience descriptions, as well as a retrieval-guided generative model for synthesizing candidate waveforms from target descriptions. The dataset comprises 100 GVS waveforms and 1,526 valid free-form sensation descriptions collected from 16 participants. Semantic analysis revealed diverse motion- and force-related sensations, localized bodily sensations, and situational associations. Compared with a participant-preserving permutation baseline, descriptions elicited by the same waveform covered fewer semantic categories (8.18 vs. 9.45) and exhibited a higher dominant-category proportion (26.97% vs. 21.25%; both P < 0.001). Building on this dataset, we implemented the generative model as a retrieval-guided one-dimensional convolutional variational autoencoder. An independent behavioral study recruited 10 participants who had not contributed to the dataset collection. Performance in discriminating congruent from incongruent waveform-visual cue pairings was significantly above chance, with an accuracy of 63.33%, d-prime = 0.70, and p < 0.001. Together, these findings demonstrate the feasibility of text-conditioned GVS synthesis and support the development of GVS as a programmable modality for semantically congruent embodied feedback across interactive scenarios.

cs.HC

Lottery and Sprint Arcade: Enabling Player-Driven Game Editing with Generative AI

Large language models (LLMs) are shifting game generation from offline automation toward play-driven modification through natural language interaction. In this work, we present a play-driven game editing system that enables players to modify a retro Space Invaders - style arcade game through voice-based natural-language commands during play. Spoken instructions are interpreted by an LLM and translated into structured updates of internal configuration parameters, allowing iterative play - edit - feedback cycles in an invader-style game environment without exposing underlying system details. The game includes approximately 100 editable configuration fields controlling mechanics, visuals, interaction patterns, and audio behavior, enabling gameplay transformation through incremental parameter changes. To investigate how users experience play-driven AI-mediated editing (RQ1) and how emergent editing patterns relate to variations in player experience (RQ2), we conducted a user study combining subjective evaluations, workload measures, and log-based analysis of editing behavior. Participants were able to modify gameplay with generally positive experiences and moderate workload, and interaction outcomes did not strongly depend on prior programming experience. Editing-log analysis revealed distinct experiential tendencies: adjustments to immediately perceptible parameters were associated with higher usability, whereas edits affecting core gameplay structures were more closely associated with enjoyment. Post-session reflections further identified diverse editing strategies, including exploratory experimentation, goal-driven structural modification, and iterative parameter tuning. These findings demonstrate that voice-driven editing can support accessible, play-driven human - AI co-creation within a structured invader-style arcade game environment.

cs.HC

WhiteTesseract: Reframing the Interpretation of Cultural Heritage through XR and Conversational AI

Cultural heritage exhibitions often struggle to sustain attention and support reflective engagement. Physical exhibitions rely on fixed interpretive aids that lack adaptability to individual backgrounds or curiosity, and their effectiveness depends heavily on a visitor's Personal Context, prior knowledge, and cultural literacy. Meanwhile, digital exhibitions prioritize convenience and accessibility but risk weakening the Physical and Social Contexts that define embodied cultural experience. WhiteTesseract addresses this gap by enabling in-situ interpretation through high-resolution XR and conversational AI. The system integrates spatial intelligence via artwork recognition to allow visitors to selectively reduce environmental distractions (via diminished reality) and engage in context-aware dialogue (via large language models). The goal is to preserve the richness of the physical and social environment while providing a flexible space for personal reflection, enhancing Personal Context without compromising physical authenticity. We deployed the system in a Claude Monet exhibition and conducted a controlled user study with 26 participants. Quantitative results showed that WhiteTesseract modulation significantly increased average viewing duration from 35.3 to 98.3 seconds (p < 0.001). Analysis of 529 visitor-AI interactions revealed that 60% extended beyond factual queries to include analytical, emotional, and comparative inquiries. These findings demonstrate how XR and AI can enrich the physical exhibition experience by supporting deeper, more personalized engagement without displacing the embodied value of cultural heritage. We discuss technical and social constraints for real-world deployment and limitations of our controlled setting.

cs.HC

Acoustic Manipulation of Tangible Janus Icons on Liquid Droplets

Interfaces that couple digital information with physical matter enable computation to be expressed through tangible motion and touch, yet typically rely on embedded actuators, rigid mechanisms, or enclosed environments. Consequently, contactless manipulation and interaction with centimeter-scale tangible elements in open settings remain difficult to achieve. Here, we present PolygonWave, a solid--fluid acoustic interface that enables transport and tangible interaction by coupling airborne ultrasound with liquid-mediated support. The system employs lightweight Janus icons with asymmetric wettability: a superhydrophobic upper surface permits dry touch interaction, while a hydrophilic lower surface couples to a water droplet resting on a superhydrophobic mesh. Focused acoustic fields generated by a 256-element phased array induce lateral forces, enabling programmable motion without mechanical contact. Systematic characterization demonstrates transport of payloads up to 525 mg across variations in icon size, droplet volume, and applied load. Beyond translation, the liquid layer functions as a reconfigurable mechanical element, enabling button-like input with self-recovery and resonance-driven vibro-visual feedback, exhibiting a peak response near 22 Hz for 200 \textmu L droplets. Liquid-mediated acoustic coupling provides a unified mechanism for mechanically expressive, touch-accessible tangible interfaces bridging acoustics, soft matter physics, and physical human--computer interaction.

physics.app-ph

Event Topology-based Visual Microphone for Amplitude and Frequency Reconstruction

Accurate vibration measurement is vital for analyzing dynamic systems across science and engineering, yet noncontact methods often balance precision against practicality. Event cameras offer high-speed, low-light sensing, but existing approaches fail to recover vibration amplitude and frequency with sufficient accuracy. We present an event topology-based visual microphone that reconstructs vibrations directly from raw event streams without external illumination. By integrating the Mapper algorithm from topological data analysis with hierarchical density-based clustering, our framework captures the intrinsic structure of event data to recover both amplitude and frequency with high fidelity. Experiments demonstrate substantial improvements over prior methods and enable simultaneous recovery of multiple sound sources from a single event stream, advancing the frontier of passive, illumination-free vibration sensing.

physics.app-ph

Touching Movement: 3D Tactile Poses for Supporting Blind People in Learning Body Movements

Visual impairments create barriers to learning physical activities, since conventional training methods rely on visual demonstrations or often inadequate verbal descriptions. This research explores 3D-printed human body models to enhance movement comprehension for blind individuals. Through a participatory design approach in collaboration with a blind designer, we developed detailed 3D models representing various body movements and incorporated tactile reference elements to enhance spatial understanding. We conducted two user studies with 10 blind participants across different activities: static yoga poses and sequential calisthenic movements. The results demonstrated that 3D models significantly improved understanding speed, reduced questions for clarification, and enhanced movement accuracy compared to conventional teaching methods. Participants consistently rated 3D models higher for ease of understanding, effectiveness, and motivation.

cs.HC

Systematic Optimization of Real-Time Diffusion Model Inference on Apple M3 Ultra

While real-time image generation using diffusion models has advanced rapidly on NVIDIA GPUs, systematic optimization research on non-CUDA platforms such as Apple Silicon remains extremely limited. In this study, we conducted comprehensive optimization experiments across 10 phases targeting the Apple M3 Ultra (60-core GPU, 512 GB unified memory) with the goal of achieving real-time camera img2img transformation. We explored a wide range of techniques including CoreML conversion, quantization, Token Merging, Neural Engine utilization, compact model exploration, frame interpolation, kNN search-based synthesis, pix2pix-turbo, optical flow frame skipping, and knowledge distillation, quantitatively evaluating the effectiveness of each approach. Ultimately, by combining CoreML conversion of the distillation-specialized model SDXS-512 with a 3-thread camera pipeline, we achieved real-time camera img2img transformation at 22.7 FPS at 512x512 resolution. The primary contribution of this work is the systematic demonstration that optimization insights established for CUDA are not necessarily effective on Apple Silicon's unified memory architecture. We reveal an optimization landscape fundamentally different from that of NVIDIA GPUs -- including the absence of speedup from quantization, the ineffectiveness of parallel inference, and the unsuitability of the Neural Engine for large-scale models -- and provide practical guidelines for diffusion model inference on Apple Silicon.

cs.LG

Reversible vertical positioning of acoustically levitated particle using a spiral reflector

Dynamic positioning in acoustic levitation typically depends on active control of the transducers phases, which necessitates complex driving electronics. While mechanically actuated reflectors offer a simpler alternative, achieving reversible transport along the vertical axis solely through mechanical actuation remains challenging. Here, we demonstrate vertical particle translation using a rotating spiral reflector with a half-wavelength pitch. With the rotation axis laterally offset relative to the acoustic focus, the spiral surface functions as a series of translating slopes. Experimental and numerical results confirm stable, bidirectional transport, yielding a vertical displacement of approximately $0.58λ$ per revolution and a maximum height of $3.18λ$, with radial confinement maintained within $0.24λ$. This approach provides a cost-effective solution for non-contact sample handling without active phase control.

physics.app-ph

OnomaCompass: A Texture Exploration Interface that Shuttles between Words and Images

Humans can finely perceive material textures, yet articulating such somatic impressions in words is a cognitive bottleneck in design ideation. We present OnomaCompass, a web-based exploration system that links sound-symbolic onomatopoeia and visual texture representations to support early-stage material discovery. Instead of requiring users to craft precise prompts for generative AI, OnomaCompass provides two coordinated latent-space maps--one for texture images and one for onomatopoeic term--built from an authored dataset of invented onomatopoeia and corresponding textures generated via Stable Diffusion. Users can navigate both spaces, trigger cross-modal highlighting, curate findings in a gallery, and preview textures applied to objects via an image-editing model. The system also supports video interpolation between selected textures and re-embedding of extracted frames to form an emergent exploration loop. We conducted a within-subjects study with 11 participants comparing OnomaCompass to a prompt-based image-generation workflow using Gemini 2.5 Flash Image ("Nano Banana"). OnomaCompass significantly reduced workload (NASA-TLX overall, mental demand, effort, and frustration; p < .05) and increased hedonic user experience (UEQ), while usability (SUS) favored the baseline. Qualitative findings indicate that OnomaCompass helps users externalize vague sensory expectations and promotes serendipitous discovery, but also reveals interaction challenges in spatial navigation. Overall, leveraging sound symbolism as a lightweight cue offers a complementary approach to Kansei-driven material ideation beyond prompt-centric generation.

cs.HC

Knowing Ourselves Through Others: Reflecting with AI in Digital Human Debates

LLMs can act as an impartial other, drawing on vast knowledge, or as personalized self-reflecting user prompts. These personalized LLMs, or Digital Humans, occupy an intermediate position between self and other. This research explores the dynamic of self and other mediated by these Digital Humans. Using a Research Through Design approach, nine junior and senior high school students, working in teams, designed Digital Humans and had them debate. Each team built a unique Digital Human using prompt engineering and RAG, then observed their autonomous debates. Findings from generative AI literacy tests, interviews, and log analysis revealed that participants deepened their understanding of AI's capabilities. Furthermore, experiencing their own creations as others prompted a reflective attitude, enabling them to objectively view their own cognition and values. We propose "Reflecting with AI" - using AI to re-examine the self - as a new generative AI literacy, complementing the conventional understanding, applying, criticism and ethics.

cs.HC

Digital Nature Revisited: A Ten-Year Synthesis of Art, Technology, and the Evolution of "Nature": Reimagining Post-Truth Ecologies Through Art, Algorithm, and Animism

This paper critically re-examines "Digital Nature," a concept that has proliferated across various domains over the last ten years. By "Digital Nature," we refer to an evolving view of nature as a dynamic process of circulating computation and matter, one that extends into the realms of AI, XR, indigenous perspectives, and post-human theory. Despite its popularity, "Digital Nature" remains ambiguously defined. This paper provides a genealogical and philosophical survey of how the idea has emerged, diverged, and overlapped in media art, bio-art, and generative art, alongside relevant Eastern, Islamic, and indigenous worldviews. We then introduce a multi-axis framework (from real/virtual to anthropocentric/object-oriented, with sub-axes of enchantment and materialization), illustrating how digital technologies have reconceptualized the question "What is nature?" in unexpected ways. Finally, we discuss how the field might evolve, particularly through the lens of large language models, AGI, and "supernatural reality," while highlighting the ethical and political pitfalls of techno-occultism. Our ultimate goal is to re-situate "Digital Nature" as both an intellectual frontier and a collaborative platform that invites continuous dialogue between art, science, technology, and cultural philosophies.

cs.HC

Suzume-chan: Your Personal Navigator as an Embodied Information Hub

Access to expert knowledge often requires real-time human communication. Digital tools improve access to information but rarely create the sense of connection needed for deep understanding. This study addresses this issue using Social Presence Theory, which explains how a feeling of "being together" enhances communication. An "Embodied Information Hub" is proposed as a new way to share knowledge through physical and conversational interaction. The prototype, Suzume-chan, is a small, soft AI agent running locally with a language model and retrieval-augmented generation (RAG). It learns from spoken explanations and responds through dialogue, reducing psychological distance and making knowledge sharing warmer and more human-centered.

cs.AI

Instant Skinned Gaussian Avatars for Web, Mobile and VR Applications

We present Instant Skinned Gaussian Avatars, a real-time and cross-platform 3D avatar system. Many approaches have been proposed to animate Gaussian Splatting, but they often require camera arrays, long preprocessing times, or high-end GPUs. Some methods attempt to convert Gaussian Splatting into mesh-based representations, achieving lightweight performance but sacrificing visual fidelity. In contrast, our system efficiently animates Gaussian Splatting by leveraging parallel splat-wise processing to dynamically follow the underlying skinned mesh in real time while preserving high visual fidelity. From smartphone-based 3D scanning to on-device preprocessing, the entire process takes just around five minutes, with the avatar generation step itself completed in only about 30 seconds. Our system enables users to instantly transform their real-world appearance into a 3D avatar, making it ideal for seamless integration with social media and metaverse applications. Website: https://gaussian-vrm.github.io/

cs.CG

Photographic Conviviality: A Synchronic and Symbiotic Photographic Experience through a Body Paint Workshop

This study explores "Photo Tattooing," merging photography and body ornamentation, and introduces the concept of "Photographic Conviviality." Using our instant camera that prints images onto mesh screens for immediate body art, we examine how this integration affects personal expression and challenges traditional photography. Workshops revealed that this fusion redefines photography's role, fostering intimacy and shared experiences, and opens new avenues for self-expression by transforming static images into dynamic, corporeal experiences.

cs.HC

null2: Boundary-Dissolving Bodies and Architecture towards Digital Nature

This paper presents a case study of the thematic pavilion null2 at Expo 2025 Osaka-Kansai, contrasting with the static Jomon motifs of Taro Okamoto's Tower of the Sun from Expo 1970. The study discusses Yayoi-inspired mirror motifs and dynamically transforming interactive spatial configuration of null2, where visitors become integrated as experiential content. The shift from static representation to a new ontological and aesthetic model, characterized by the visitor's body merging in real-time with architectural space at installation scale, is analyzed. Referencing the philosophical context of Expo 1970 theme 'Progress and Harmony for Mankind,' this research reconsiders the worldview articulated by null2 in Expo 2025, in which computation is naturalized and ubiquitous, through its intersection with Eastern philosophical traditions. It investigates how immersive experiences within the pavilion, grounded in the philosophical framework of Digital Nature, reinterpret traditional spatial and structural motifs of the tea room, positioning them within contemporary digital art discourse. The aim is to contextualize and document null2 as an important contemporary case study from Expo practices, considering the historical and social background in Japan from the 19th to 21st century, during which world expositions served as pivotal points for the birth of modern Japanese concept of 'fine art,' symbolic milestones of economic development, and key moments in urban and media culture formation. Furthermore, this paper academically organizes architectural techniques, computer graphics methodologies, media art practices, and theoretical backgrounds utilized in null2, highlighting the scholarly significance of preserving these as an archival document for future generations.

cs.HC

Dynamic Caustics by Ultrasonically Modulated Liquid Surface

This paper presents a method for generating dynamic caustic patterns by utilising dual-optimised holographic fields with Phased Array Transducer (PAT). Building on previous research in static caustic optimisation and ultrasonic manipulation, this approach employs computational techniques to dynamically shape fluid surfaces, thereby creating controllable and real-time caustic images. The system employs a Digital Twin framework, which enables iterative feedback and refinement, thereby improving the accuracy and quality of the caustic patterns produced. This paper extends the foundational work in caustic generation by integrating liquid surfaces as refractive media. This concept has previously been explored in simulations but not fully realised in practical applications. The utilisation of ultrasound to directly manipulate these surfaces enables the generation of dynamic caustics with a high degree of flexibility. The Digital Twin approach further enhances this process by allowing for precise adjustments and optimisation based on real-time feedback. Experimental results demonstrate the technique's capacity to generate continuous animations and complex caustic patterns at high frequencies. Although there are limitations in contrast and resolution compared to solid-surface methods, this approach offers advantages in terms of real-time adaptability and scalability. This technique has the potential to be applied in a number of areas, including interactive displays, artistic installations and educational tools. This research builds upon the work of previous researchers in the fields of caustics optimisation, ultrasonic manipulation, and computational displays. Future research will concentrate on enhancing the resolution and intricacy of the generated patterns.

cs.GR

Insect-Computer Hybrid Speaker: Speaker using Chirp of the Cicada Controlled by Electrical Muscle Stimulation

We propose "Insect-Computer Hybrid Speaker", which enables us to make musics made from combinations of computer and insects. Lots of studies have proposed methods and interfaces for controlling insects and obtaining feedback. However, there have been less research on the use of insects for interaction with third parties. In this paper, we propose a method in which cicadas are used as speakers triggered by using Electrical Muscle Stimulation (EMS). We explored and investigated the suitable waveform of chirp to be controlled, the appropriate voltage range, and the maximum pitch at which cicadas can chirp.

cs.HC