SearcharxivSearch

arXiv subjects

Ryo Watanabe

Publications and source records attributed to Ryo Watanabe.

13 recordsLinked to original sources

When Absolute State Fails: Evaluating Proprioceptive Encodings for Robust Manipulation

As end-to-end robotic policies are progressively deployed in the real world to solve real tasks, they face a gap between the training and inference conditions. Scaling the amount and diversity of the training data has shown some success in improving zero-shot generalization, yet robots still fail when faced with new, unseen test conditions. For instance, while robots with fixed frames of reference are common, those with moving frames pose a greater challenge for deployment. To address this specific instance of the issue, we present a study of strategies for encoding the robot's proprioceptive state to improve both in- and out-of-distribution performance at test time. Through a systematic study of joint representations, we find that a simple episode-wise relative frame provides the best trade-off between task performance and robustness, outperforming the baselines in extensive real-robot experiments conducted in a realistic test environment. The results suggest a practical path to leveraging data collected by robots with varying frames of reference and deployment to unseen test configurations.

cs.RO

Tensor network surrogate models for variational quantum computation

We adopt a two-dimensional tensor-network (TN) ansatz to simulate variational quantum algorithms on two-dimensional qubit architectures, demonstrating its capability to accurately simulate deep circuits through the Quantum Approximate Optimization Algorithm (QAOA) applied to Ising spin-glass problems on heavy-hexagonal and square lattices. For heavy-hexagonal problems with up to three-body interactions, parameters trained on small instances and transferred to systems an order of magnitude larger improve the sampled energy distribution only up to intermediate depths, indicating a fundamental limit of parameter concentration as a transfer strategy. By extending the training itself with TN simulations on larger system sizes, we avoid local minima and obtain lower-energy samples. Analyses of entanglement growth and importance sampling show that the simulation remains classically feasible with moderate bond dimension. We find that parameter concentration also persists on square lattices, albeit at substantially higher computational cost to perform reliable sampling. Overall, our TN framework not only provides an efficient and controlled framework for benchmarking variational quantum algorithms on two-dimensional lattices, but also serves as an effective surrogate model for training variational algorithms.

quant-ph

Quantum-Inspired Algorithm for Classical Spin Hamiltonians Based on Matrix Product Operators

We propose a tensor-network (TN) approach for solving classical optimization problems that is inspired by spectral filtering and sampling on quantum states. We first shift and scale an Ising Hamiltonian of the cost function so that all eigenvalues become non-negative and the ground states correspond to the the largest eigenvalues, which are then amplified by power iteration. We represent the transformed Hamiltonian as a matrix product operator (MPO) and form an immense power of this object via truncated MPO-MPO contractions, embedding the resulting operator into a matrix product state for sampling in the computational basis. In contrast to the density-matrix renormalization group, our approach provides a straightforward route to systematic improvement by increasing the bond dimension and is better at avoiding local minima. We also study the performance of this power method in the context of a higher-order Ising Hamiltonian on a heavy-hexagonal lattice, making a comparison with simulated annealing. These results highlight the potential of quantum-inspired algorithms for solving optimization problems and provide a baseline for assessing and developing quantum algorithms.

quant-ph

Site-Order Optimization in the Density Matrix Renormalization Group via Multi-Site Rearrangement

In the approaches based on matrix-product states (MPSs), such as the density-matrix renormalization group (DMRG) method, the ordering of the sites crucially affects the computational accuracy. We investigate the performance of an algorithm that searches for the optimal site order by iterative local site rearrangement. We improve the algorithm by expanding the range of site rearrangement and apply it to a one-dimensional quantum Heisenberg model with random site permutation. The results indicate that increasing the range of the site rearrangement significantly improves the computational accuracy of the DMRG method. In particular, increasing the rearrangement range from two to three sites reduces the average relative error in the ground-state energy by 65% to 94% in the cases we tested. We also discuss the computational cost of the algorithm and its application as a preprocessing for MPS-based calculations.

cond-mat.stat-mech

Improving Visual Recommendation on E-commerce Platforms Using Vision-Language Models

On large-scale e-commerce platforms with tens of millions of active monthly users, recommending visually similar products is essential for enabling users to efficiently discover items that align with their preferences. This study presents the application of a vision-language model (VLM) -- which has demonstrated strong performance in image recognition and image-text retrieval tasks -- to product recommendations on Mercari, a major consumer-to-consumer marketplace used by more than 20 million monthly users in Japan. Specifically, we fine-tuned SigLIP, a VLM employing a sigmoid-based contrastive loss, using one million product image-title pairs from Mercari collected over a three-month period, and developed an image encoder for generating item embeddings used in the recommendation system. Our evaluation comprised an offline analysis of historical interaction logs and an online A/B test in a production environment. In offline analysis, the model achieved a 9.1% improvement in nDCG@5 compared with the baseline. In the online A/B test, the click-through rate improved by 50% whereas the conversion rate improved by 14% compared with the existing model. These results demonstrate the effectiveness of VLM-based encoders for e-commerce product recommendations and provide practical insights into the development of visual similarity-based recommendation systems.

cs.IR

FTACT: Force Torque aware Action Chunking Transformer for Pick-and-Reorient Bottle Task

Manipulator robots are increasingly being deployed in retail environments, yet contact rich edge cases still trigger costly human teleoperation. A prominent example is upright lying beverage bottles, where purely visual cues are often insufficient to resolve subtle contact events required for precise manipulation. We present a multimodal Imitation Learning policy that augments the Action Chunking Transformer with force and torque sensing, enabling end-to-end learning over images, joint states, and forces and torques. Deployed on Ghost, single-arm platform by Telexistence Inc, our approach improves Pick-and-Reorient bottle task by detecting and exploiting contact transitions during pressing and placement. Hardware experiments demonstrate greater task success compared to baseline matching the observation space of ACT as an ablation and experiments indicate that force and torque signals are beneficial in the press and place phases where visual observability is limited, supporting the use of interaction forces as a complementary modality for contact rich skills. The results suggest a practical path to scaling retail manipulation by combining modern imitation learning architectures with lightweight force and torque sensing.

cs.RO

TTNOpt: Tree tensor network package for high-rank tensor compression

We have developed TTNOpt, a software package that utilizes tree tensor networks (TTNs) for quantum spin systems and high-dimensional data analysis. TTNOpt provides efficient and powerful TTN computations by locally optimizing the network structure, guided by the entanglement pattern of the target tensors. For quantum spin systems, TTNOpt searches for the ground state of Hamiltonians with bilinear spin interactions and magnetic fields, and computes physical properties of these states, including the variational energy, bipartite entanglement entropy (EE), single-site expectation values, and two-site correlation functions. Additionally, TTNOpt can target the lowest-energy state within a specified subspace, provided that the Hamiltonian conserves total magnetization. For high-dimensional data analysis, TTNOpt factorizes complex tensors into TTN states that maximize fidelity to the original tensors by optimizing the tensors and the network. When a TTN is provided as input, TTNOpt reconstructs the network based on the EE without referencing the fidelity of the original state. We present three demonstrations of TTNOpt: (1) Ground-state search for the hierarchical chain model with a system size of $256$. The entanglement patterns of the ground state manifest themselves in a tree structure, and TTNOpt successfully identifies the tree. (2) Factorization of a quantic tensor of the $2^{24}$ dimensions representing a three-variable function where each variant has a weak bit-wise correlation. The optimized TTN shows that its structure isolates the variables from each other. (3) Reconstruction of the matrix product network representing a $16$-variable normal distribution characterized by a tree-like correlation structure. TTNOpt can reveal hidden correlation structures of the covariance matrix.

quant-ph

DFM: Deep Fourier Mimic for Expressive Dance Motion Learning

As entertainment robots gain popularity, the demand for natural and expressive motion, particularly in dancing, continues to rise. Traditionally, dancing motions have been manually designed by artists, a process that is both labor-intensive and restricted to simple motion playback, lacking the flexibility to incorporate additional tasks such as locomotion or gaze control during dancing. To overcome these challenges, we introduce Deep Fourier Mimic (DFM), a novel method that combines advanced motion representation with Reinforcement Learning (RL) to enable smooth transitions between motions while concurrently managing auxiliary tasks during dance sequences. While previous frequency domain based motion representations have successfully encoded dance motions into latent parameters, they often impose overly rigid periodic assumptions at the local level, resulting in reduced tracking accuracy and motion expressiveness, which is a critical aspect for entertainment robots. By relaxing these locally periodic constraints, our approach not only enhances tracking precision but also facilitates smooth transitions between different motions. Furthermore, the learned RL policy that supports simultaneous base activities, such as locomotion and gaze control, allows entertainment robots to engage more dynamically and interactively with users rather than merely replaying static, pre-designed dance routines.

cs.RO

Learning Quiet Walking for a Small Home Robot

As home robotics gains traction, robots are increasingly integrated into households, offering companionship and assistance. Quadruped robots, particularly those resembling dogs, have emerged as popular alternatives for traditional pets. However, user feedback highlights concerns about the noise these robots generate during walking at home, particularly the loud footstep sound. To address this issue, we propose a sim-to-real based reinforcement learning (RL) approach to minimize the foot contact velocity highly related to the footstep sound. Our framework incorporates three key elements: learning varying PD gains to actively dampen and stiffen each joint, utilizing foot contact sensors, and employing curriculum learning to gradually enforce penalties on foot contact velocity. Experiments demonstrate that our learned policy achieves superior quietness compared to a RL baseline and the carefully handcrafted Sony commercial controllers. Furthermore, the trade-off between robustness and quietness is shown. This research contributes to developing quieter and more user-friendly robotic companions in home environments.

cs.RO

Automatic Structural Search of Tensor Network States including Entanglement Renormalization

Tensor network (TN) states, including entanglement renormalization (ER), can encompass a wider variety of entangled states. When the entanglement structure of the quantum state of interest is non-uniform in real space, accurately representing the state with a limited number of degrees of freedom hinges on appropriately configuring the TN to align with the entanglement pattern. However, a proposal has yet to show a structural search of ER due to its high computational cost and the lack of flexibility in its algorithm. In this study, we conducted an optimal structural search of TN, including ER, based on the reconstruction of their local structures with respect to variational energy. Firstly, we demonstrated that our algorithm for the spin-$1/2$ tetramer singlets model could calculate exact ground energy using the multi-scale entanglement renormalization ansatz (MERA) structure as an initial TN structure. Subsequently, we applied our algorithm to the random XY models with the two initial structures: MERA and the suitable structure underlying the strong disordered renormalization group. We found that, in both cases, our algorithm achieves improvements in variational energy, fidelity, and entanglement entropy. The degree of improvement in these quantities is superior in the latter case compared to the former, suggesting that utilizing an existing TN design method as a preprocessing step is important for maximizing our algorithm's performance.

quant-ph

Fully Spiking Denoising Diffusion Implicit Models

Spiking neural networks (SNNs) have garnered considerable attention owing to their ability to run on neuromorphic devices with super-high speeds and remarkable energy efficiencies. SNNs can be used in conventional neural network-based time- and energy-consuming applications. However, research on generative models within SNNs remains limited, despite their advantages. In particular, diffusion models are a powerful class of generative models, whose image generation quality surpass that of the other generative models, such as GANs. However, diffusion models are characterized by high computational costs and long inference times owing to their iterative denoising feature. Therefore, we propose a novel approach fully spiking denoising diffusion implicit model (FSDDIM) to construct a diffusion model within SNNs and leverage the high speed and low energy consumption features of SNNs via synaptic current learning (SCL). SCL fills the gap in that diffusion models use a neural network to estimate real-valued parameters of a predefined probabilistic distribution, whereas SNNs output binary spike trains. The SCL enables us to complete the entire generative process of diffusion models exclusively using SNNs. We demonstrate that the proposed method outperforms the state-of-the-art fully spiking generative model.

cs.CV

A comprehensive survey on quantum computer usage: How many qubits are employed for what purposes?

Quantum computers (QCs), which work based on the law of quantum mechanics, are expected to be faster than classical computers in several computational tasks such as prime factoring and simulation of quantum many-body systems. In the last decade, research and development of QCs have rapidly advanced. Now hundreds of physical qubits are at our disposal, and one can find several remarkable experiments actually outperforming the classical computer in a specific computational task. On the other hand, it is unclear what the typical usages of the QCs are. Here we conduct an extensive survey on the papers that are posted in the quant-ph section in arXiv and claim to have used QCs in their abstracts. To understand the current situation of the research and development of the QCs, we evaluated the descriptive statistics about the papers, including the number of qubits employed, QPU vendors, application domains and so on. Our survey shows that the annual number of publications is increasing, and the typical number of qubits employed is about six to ten, growing along with the increase in the quantum volume (QV). Most of the preprints are devoted to applications such as quantum machine learning, condensed matter physics, and quantum chemistry, while quantum error correction and quantum noise mitigation use more qubits than the other topics. These imply that the increase in QV is fundamentally relevant, and more experiments for quantum error correction, and noise mitigation using shallow circuits with more qubits will take place.

quant-ph

Variational quantum eigensolver with embedded entanglement using a tensor-network ansatz

In this paper, we introduce a tensor network (TN) scheme into the entanglement augmentation process of the synergistic optimization framework by Rudolph et al. [arXiv:2208.13673] to build its process systematically for inhomogeneous systems. Our synergistic approach first embeds the variational optimal solution of the TN state with the entropic area law, which can be perfectly optimized in conventional (classical) computers, in a quantum variational circuit ansatz inspired by the TN state with the entropic volume law. Next, the framework performs a variational quantum eigensolver (VQE) process with embedded states as the initial state. We applied the synergistic to the ground-state analysis of the all-to-all coupled random transverse-field Ising, XYZ, Heisenberg model, employing the binary multiscale entanglement renormalization ansatz (MERA) state and branching MERA states as TN states with entropic area law and volume law, respectively. We then show that the synergistic accelerates VQE calculations in the three models without an initial parameter guess of the branching-MERA-inspired ansatz and can avoid a local solution trapped by a standard VQE with the ansatz in the Ising model. The improvement of optimizers for MERA in all-to-all coupled inhomogeneous systems, enhancement, and potential synergistic applications are also discussed.

quant-ph