SearcharxivSearch

arXiv subjects

Ryo Hanai

Publications and source records attributed to Ryo Hanai.

At least 19 recordsLinked to original sources

Non-Prehensile Throwing: A Reinforcement Learning Perspective

Robotic throwing enables fast object transport and extends a robot's reachable workspace beyond traditional pick-and-place. While prehensile (grasp-based) throwing works well for graspable items, non-prehensile (grasp-free) throwing is better suited for large, heavy, and/or deformable objects. Existing approaches rely on model-based optimization with simplified contact models (e.g., dynamic grasping) and low-dimensional trajectory parameterizations, which limit solution quality and reachable workspace. We propose a reinforcement learning approach that additionally leverages sliding and rolling contact modes and directly optimizes joint-space trajectories without analytical contact models or custom parameterizations. The Markov Decision Process (MDP) is formulated as a dynamical system that evolves the robot's joint state conditioned on the throwing target, object model, and initial configuration. Joint-jerk trajectories are planned offline at a low control rate and upsampled into smooth, high-rate velocity commands for deployment. For sim-to-real transfer, we minimize the robot-dynamics gap through minimum-jerk system identification and train uncertainty-aware policies to mitigate object-modeling errors, particularly sensitivity to dynamic friction. In simulation, the policy achieves 99% success across thousands of configurations and generalizes to unseen objects. Sensitivity analysis shows robustness to mass uncertainty but high sensitivity to dynamic friction, consistent with the sliding-based release mechanism. Deployed zero-shot on a UR5e operating near its physical limits (5 m/s end-effector velocity), our method throws diverse objects including heavy (790 g) and large (20x20x28 cm) items to targets up to 350 cm distance or 180 cm elevation, achieving a 97% real-world success rate.

cs.RO

The Embodiment Gap in Robot Foundation Models

Robot foundation models (RFMs), including vision-language-action (VLA) policies, are often discussed through a scaling view: more data, larger models, and broader benchmarks should improve generalization. In robotics, however, a model can generalize while work still remains before it can run on a robot with a particular body. The work required differs across methods and target robots, and those differences affect practical deployment. We call the gap between reusable models, representations, or data and their use in execution on the target robot the embodiment gap. This survey examines what can be reused across robot embodiments and what must still be implemented on a new robot. We place existing methods on a two-axis map that shows the type of shared structure and the stage at which adaptation is needed for execution on the target robot. We then examine recent work through three overlapping research directions: sharing semantics and perception, sharing robot data and interfaces, and learning correspondence across embodiments. We also propose a reporting framework for adaptation work that success rate alone does not reveal. The framework identifies the work that should be checked when comparing cross-embodiment learning and highlights work that remains on a new robot and questions for future study.

cs.RO

ORPA: Online Residual Policy Adaptation for Robot Manipulation Control with Human Feedback

Robotic manipulation policies trained via imitation learning, such as Action Chunking with Transformers (ACT), can achieve strong performance under ideal conditions but often remain sensitive to small execution errors and distribution shifts. Correcting these failures typically requires dataset aggregation and full-policy retraining, which is computationally expensive and unsuitable for real-time deployment. In this work, we propose Online Residual Policy Adaptation (ORPA), a framework that enables immediate, feedback-driven correction of robot actions without modifying the underlying policy parameters. ORPA augments a pretrained control policy with a lightweight, feedback-conditioned module that predicts residual adjustments directly in joint space, allowing the system to adapt its behavior at runtime. We evaluate ORPA on a set of precision-sensitive manipulation tasks using the ALOHA platform, demonstrating improvements in success rate and recovery from small perturbations compared to baseline control policies and rule-based inverse kinematics corrections.

cs.RO

Learning to Predict Contact Force Distributions from Vision Leveraging Object Geometry Priors

Based on vision and prior experience, humans can make rough physical predictions and adjust their manipulation strategies. This paper aims to endow robots with a similar ability. To collect paired data of vision and forces, we use a rigid-body simulator commonly adopted in robotics. However, unlike simulators that output noisy point forces, humans are able to make consistent predictions even in unfamiliar situations. Based on this observation, we hypothesize that predicting smooth force distributions rather than raw point forces can improve both force prediction itself and downstream task performance. To validate this hypothesis, we construct a model that predicts three-dimensional force distributions from a single RGB image of piled daily objects. The target distribution is generated by applying statistical smoothing to point forces obtained from the simulator. Moreover, by incorporating object geometry into the smoothing process, we aim to account for variations in contact states and achieve more consistent vision-based predictions. We conduct extensive evaluations in both simulation and real environments. Results show that our approach improves prediction accuracy, enhances downstream task performance through smoothing, and further benefits from geometry-guided smoothing. Remarkably, the trained model generalizes effectively to real-world scenes despite being trained solely in simulation.

cs.RO

GuidedAttention: Interpretable and Correctable Visual Attention for OOD-Robust Robot Manipulation via Imitation Learning

End-to-end visuomotor policies provide little opportunity for humans to understand or correct the policy's visual attention. We propose GuidedAttention, a visuomotor imitation learning framework that introduces interpretable and correctable visual attention as an explicit intermediate representation. Task-relevant attention keypoints are predicted from camera images and condition a diffusion-based action policy. Users can inspect and optionally correct selected keypoints once at rollout initialization, after which the corrected attention is automatically propagated throughout execution by a tracking module. Experiments in simulation and the real world demonstrate that GuidedAttention consistently improves robot manipulation performance, particularly under positional and appearance out-of-distribution (OOD) conditions. https://mmurooka.github.io/guided-attention-project-page

cs.RO

Theory of Andreev and shot noise spectroscopy for topological superconductors probed by $s$-wave superconducting tips

Scanning tunneling microscopy (STM) and spectroscopy (STS) with $s$-wave superconducting tips has been widely applied to probe exotic superconductors, but its potential for investigating topological superconductors remains unclear. In junctions between an $s$-wave superconductor and a topological superconductor, the dominant tunneling process is Andreev reflection, in which Cooper pairs from the $s$-wave superconductor tunnel as particle--hole excitations into the surface state of the topological superconductor. In this work, we theoretically investigate the fundamental properties of Andreev and shot noise spectroscopy on topological superconductors, focusing on the $dI/dV$ characteristics and current noise. We develop a real-time description of an effective tunneling action incorporating Andreev reflection processes in the Keldysh formalism and derive analytical expressions for the Andreev reflection current and the associated current noise. Furthermore, we perform numerical simulations for representative topological superconductors and provide a catalog of $dI/dV$ spectra and the Fano factor. Our results establish guidelines for probing topological superconductivity using STM with $s$-wave superconducting tips, and provide theoretical benchmarks for future STS experiments.

cond-mat.supr-con

A Flexible Field-Based Policy Learning Framework for Diverse Robotic Systems and Sensors

We present a cross robot visuomotor learning framework that integrates diffusion policy based control with 3D semantic scene representations from D3Fields to enable category level generalization in manipulation. Its modular design supports diverse robot camera configurations including UR5 arms with Microsoft Azure Kinect arrays and bimanual manipulators with Intel RealSense sensors through a low latency control stack and intuitive teleoperation. A unified configuration layer enables seamless switching between setups for flexible data collection training and evaluation. In a grasp and lift block task the framework achieved an 80 percent success rate after only 100 demonstration episodes demonstrating robust skill transfer between platforms and sensing modalities. This design paves the way for scalable real world studies in cross robotic generalization.

cs.RO

Quantum non-Markovian Hatano-Nelson model

While considering non-Hermitian Hamiltonians arising in the presence of dissipation, in most cases, the dissipation is taken to be frequency independent. However, this idealization may not always be applicable in experimental settings, where dissipation can be frequency-dependent. Such frequency-dependent dissipation leads to non-Markovian behavior. In this work, we demonstrate how a non-Markovian generalization of the Hatano-Nelson model, a paradigmatic non-Hermitian system with nonreciprocal hopping, arises microscopically in a quasi-one-dimensional dissipative lattice. This is achieved using non-equilibrium Green's functions without requiring any approximation like weak system-bath coupling or a time-scale separation, which would have been necessary for a Markovian treatment. The resulting effective system exhibits nonreciprocal hopping, as well as uniform dissipation, both of which are frequency-dependent. This holds for both bosonic and fermionic settings. We find solely non-Markovian nonreciprocal features like unidirectional frequency blocking in bosonic setting, and a non-equilibrium dissipative quantum phase transition in fermionic setting, that cannot be captured in a Markovian theory, nor have any analog in reciprocal systems. Our results lay the groundwork for describing and engineering non-Markovian nonreciprocal quantum lattices.

quant-ph

NeuralMeshing: Complete Object Mesh Extraction from Casual Captures

How can we extract complete geometric models of objects that we encounter in our daily life, without having access to commercial 3D scanners? In this paper we present an automated system for generating geometric models of objects from two or more videos. Our system requires the specification of one known point in at least one frame of each video, which can be automatically determined using a fiducial marker such as a checkerboard or Augmented Reality (AR) marker. The remaining frames are automatically positioned in world space by using Structure-from-Motion techniques. By using multiple videos and merging results, a complete object mesh can be generated, without having to rely on hole filling. Code for our system is available from https://github.com/FlorisE/NeuralMeshing.

cs.CV

Visual Prompting for Robotic Manipulation with Annotation-Guided Pick-and-Place Using ACT

Robotic pick-and-place tasks in convenience stores pose challenges due to dense object arrangements, occlusions, and variations in object properties such as color, shape, size, and texture. These factors complicate trajectory planning and grasping. This paper introduces a perception-action pipeline leveraging annotation-guided visual prompting, where bounding box annotations identify both pickable objects and placement locations, providing structured spatial guidance. Instead of traditional step-by-step planning, we employ Action Chunking with Transformers (ACT) as an imitation learning algorithm, enabling the robotic arm to predict chunked action sequences from human demonstrations. This facilitates smooth, adaptive, and data-driven pick-and-place operations. We evaluate our system based on success rate and visual analysis of grasping behavior, demonstrating improved grasp accuracy and adaptability in retail environments.

cs.RO

Generalized non-reciprocal phase transitions in multipopulation systems

Non-reciprocal interactions are prevalent in various complex systems leading to phenomena that cannot be described by traditional equilibrium statistical physics. Although non-reciprocally interacting systems composed of two populations have been closely studied, the physics of non-reciprocal systems with a general number of populations is not well explored despite the relevance to biological systems, active matter, and driven-dissipative quantum materials. In this work, we investigate the generic features of the phases and phase transitions and emerge in $O(2)$ symmetric many-body systems with multiple non-reciprocally coupled populations, applicable to microscopic models such as networks of oscillators, flocking models, and more generally systems where each agent has a phase variable. Using symmetry and topology of the possible orbits, we systematically show that a rich variety of time-dependent phases and phase transitions arise. Examples include multipopulation chiral phases that are distinct from their two-population counterparts that emerge via a phase transition characterized by critical exceptional points, as well as limit cycle saddle-node bifurcation and Hopf bifurcation. Interestingly, we find a phase transition that dynamically restores the $\mathbb{Z}_2$ symmetry occurs via a homoclinic orbit bifurcation, where the two $\mathbb{Z}_2$ broken orbits merge at the phase transition point, providing a general route to homoclinic chaos in the order parameter dynamics for $N\geq4$ populations. Our framework provides general principles for understanding non-equilibrium heterogeneous systems and guides experimental exploration into such systems.

cond-mat.soft

Learning Bimanual Manipulation via Action Chunking and Inter-Arm Coordination with Transformers

Robots that can operate autonomously in a human living environment are necessary to have the ability to handle various tasks flexibly. One crucial element is coordinated bimanual movements that enable functions that are difficult to perform with one hand alone. In recent years, learning-based models that focus on the possibilities of bimanual movements have been proposed. However, the high degree of freedom of the robot makes it challenging to reason about control, and the left and right robot arms need to adjust their actions depending on the situation, making it difficult to realize more dexterous tasks. To address the issue, we focus on coordination and efficiency between both arms, particularly for synchronized actions. Therefore, we propose a novel imitation learning architecture that predicts cooperative actions. We differentiate the architecture for both arms and add an intermediate encoder layer, Inter-Arm Coordinated transformer Encoder (IACE), that facilitates synchronization and temporal alignment to ensure smooth and coordinated actions. To verify the effectiveness of our architectures, we perform distinctive bimanual tasks. The experimental results showed that our model demonstrated a high success rate for comparison and suggested a suitable architecture for the policy learning of bimanual manipulation.

cs.RO

Critical scaling in one-dimensional non-reciprocal matter

Unveiling universal non-equilibrium scaling laws has been a central theme in modern statistical physics, with recent attention increasingly directed toward non-equilibrium phases that exhibit rich dynamical phenomena. A striking example arises in non-reciprocal systems, where asymmetric interactions between components lead to inherently dynamic phases and unconventional criticality near a critical exceptional point (CEP), where the criticality arises from the coalescence of collective modes with an existing Nambu-Goldstone mode. However, the scaling behavior that emerges in this system with full consideration of many-body effects and stochastic noise remains largely elusive. Here, we establish a dynamical scaling law in a generic one-dimensional (1D) stochastic non-reciprocal $O(2)$-symmetric system. Through large-scale simulations, we uncover a new non-equilibrium scaling in the vicinity of the CEP, distinct from any previously known equilibrium or non-equilibrium universality classes. We report an anomalously large roughening exponent $\alpha_{\rm CEP}=1.35(5)$, which is to be compared with those of simple diffusion $\alpha_{\rm EW}=1/2$. In regimes where the system breaks into domains with opposite chirality and spatiotemporal vortices inevitably emerge, we find that fluctuations are strongly suppressed, leading to a logarithmic scaling as a function of system size $L$ that manifests a short-range correlation. This work elucidates the beyond-mean-field dynamics of non-reciprocal matter, thereby shedding light on the exploration of criticality in non-reciprocal phase transition across diverse physical contexts, from active matter and driven quantum systems to biological pattern formation and non-Hermitian physics.

cond-mat.stat-mech

Attention-Guided Integration of CLIP and SAM for Precise Object Masking in Robotic Manipulation

This paper introduces a novel pipeline to enhance the precision of object masking for robotic manipulation within the specific domain of masking products in convenience stores. The approach integrates two advanced AI models, CLIP and SAM, focusing on their synergistic combination and the effective use of multimodal data (image and text). Emphasis is placed on utilizing gradient-based attention mechanisms and customized datasets to fine-tune performance. While CLIP, SAM, and Grad- CAM are established components, their integration within this structured pipeline represents a significant contribution to the field. The resulting segmented masks, generated through this combined approach, can be effectively utilized as inputs for robotic systems, enabling more precise and adaptive object manipulation in the context of convenience store products.

cs.RO

Phase Transitions in Nonreciprocal Driven-Dissipative Condensates

We investigate the influence of boundaries and spatial nonreciprocity on nonequilibrium driven-dissipative phase transitions. We focus on a one-dimensional lattice of nonlinear bosons described by a Lindblad master equation, where the interplay between coherent and incoherent dynamics generates nonreciprocal interactions between sites. Using a mean-field approach, we analyze the phase diagram under both periodic and open boundary conditions. For periodic boundaries, the system always forms a condensate at nonzero momentum and frequency, resulting in a time-dependent traveling wave pattern. In contrast, open boundaries reveal a far richer phase diagram, featuring multiple static and dynamical phases, as well as exotic phase transitions, including the spontaneous breaking of particle-hole symmetry associated with a critical exceptional point and phases with distinct bulk and edge behavior. Our model does not require post-selection and is experimentally realizable in platforms such as superconducting circuits.

quant-ph

Non-reciprocal interactions drive emergent chiral crystallites

We study a new type of 2D active material that exhibits macroscopic phases with two emergent broken symmetries: self-propelled achiral particles that form dense hexatic clusters, which spontaneously rotate. We experimentally realise active colloids that self-organise into both polar and hexatic crystallites, exhibiting exotic emergent phenomena. This is accompanied by a field theory of coupled order parameters formulated on symmetry principles, including non-reciprocity, to capture the non-equilibrium dynamics. We find that the presence of two interacting broken symmetry fields leads to the emergence of novel chiral phases built from (2D) achiral active colloids (here Quincke rollers). These phases are characterised by the presence of both clockwise and counterclockwise rotating clusters. We thus show that spontaneous rotation can emerge in non-equilibrium systems, even when the building blocks are achiral, due to non-reciprocally coupled broken symmetries. This interplay leads to self-organized stirring through counter-rotating vortices in confined colloidal systems, with cluster size controlled by external electric fields.

cond-mat.soft

SuctionPrompt: Visual-assisted Robotic Picking with a Suction Cup Using Vision-Language Models and Facile Hardware Design

The development of large language models and vision-language models (VLMs) has resulted in the increasing use of robotic systems in various fields. However, the effective integration of these models into real-world robotic tasks is a key challenge. We developed a versatile robotic system called SuctionPrompt that utilizes prompting techniques of VLMs combined with 3D detections to perform product-picking tasks in diverse and dynamic environments. Our method highlights the importance of integrating 3D spatial information with adaptive action planning to enable robots to approach and manipulate objects in novel environments. In the validation experiments, the system accurately selected suction points 75.4%, and achieved a 65.0% success rate in picking common items. This study highlights the effectiveness of VLMs in robotic manipulation tasks, even with simple 3D processing.

cs.RO

Visual Imitation Learning of Non-Prehensile Manipulation Tasks with Dynamics-Supervised Models

Unlike quasi-static robotic manipulation tasks like pick-and-place, dynamic tasks such as non-prehensile manipulation pose greater challenges, especially for vision-based control. Successful control requires the extraction of features relevant to the target task. In visual imitation learning settings, these features can be learnt by backpropagating the policy loss through the vision backbone. Yet, this approach tends to learn task-specific features with limited generalizability. Alternatively, learning world models can realize more generalizable vision backbones. Utilizing the learnt features, task-specific policies are subsequently trained. Commonly, these models are trained solely to predict the next RGB state from the current state and action taken. But only-RGB prediction might not fully-capture the task-relevant dynamics. In this work, we hypothesize that direct supervision of target dynamic states (Dynamics Mapping) can learn better dynamics-informed world models. Beside the next RGB reconstruction, the world model is also trained to directly predict position, velocity, and acceleration of environment rigid bodies. To verify our hypothesis, we designed a non-prehensile 2D environment tailored to two tasks: "Balance-Reaching" and "Bin-Dropping". When trained on the first task, dynamics mapping enhanced the task performance under different training configurations (Decoupled, Joint, End-to-End) and policy architectures (Feedforward, Recurrent). Notably, its most significant impact was for world model pretraining boosting the success rate from 21% to 85%. Although frozen dynamics-informed world models could generalize well to a task with in-domain dynamics, but poorly to a one with out-of-domain dynamics.

cs.RO