SearcharxivSearch

arXiv subjects

Jiaping Wu

Publications and source records attributed to Jiaping Wu.

4 recordsLinked to original sources

CAAT: Contact-Aware Attention Scaling and Tactile Masking for Data-Efficient Contact-Rich Manipulation

In contact-rich manipulation, visual observations primarily guide motion in free space, whereas tactile observations become particularly informative during contact. However, standard Transformer-based visuo-tactile policies typically rely on either token concatenation or learnable gating. These approaches lack explicit contact-aware priors, making it difficult to efficiently learn effective cross-modal representations from demonstrations. To address this limitation, we propose CAAT, a lightweight contact-aware framework that explicitly incorporates contact priors through attention scaling and dynamic tactile masking. Specifically, CAAT emphasizes visual information before contact and tactile information during contact. It also suppresses static background tokens by comparing the current tactile observation with a non-contact reference. CAAT can be integrated into commonly used Transformer-based policies without modifying their action decoders. In simulation, integrating CAAT with ACT improves the average success rate by 18.0 percentage points over direct visuo-tactile fusion and by 10.0 percentage points over gated fusion. In real-world experiments using a visuo-tactile UMI platform, CAAT achieves an average success rate of 60.0% across ACT, Diffusion Policy, and $\pi_0$, outperforming the strongest baseline by an average of 21.1 percentage points. These results demonstrate that explicit contact priors and dynamic tactile masking are effective in improving visuo-tactile policy learning and task performance of diverse policy architectures. https://mrjiangjm.github.io/caat/

cs.RO

HT-Bench: Benchmarking and Learning Dexterous Full-Hand Tactile Representations with Egocentric Vision

Establishing a universal benchmark for tactile representation learning in robotic manipulation remains challenging due to the diversity of tactile sensor designs, data formats, and robot embodiments. Rather than seeking to establish such, we explore a scalable and promising direction for future development: egocentric vision paired with full-hand tactile data. To this end, we introduce \textbf{HT-Bench}, a large-scale multi-task benchmark for dexterous full-hand tactile sensing, comprising 10M RGB frames and 7.8M tactile frames collected across 226 tasks. HT-Bench evaluates tactile representations from three key perspectives: whether they encode meaningful contact geometry, whether they can align tactile observations with visual information, and whether they generalize to unseen tasks. To assess these capabilities, HT-Bench includes four tasks: fine-grained tactile similarity retrieval, masked tactile inpainting, vision-to-tactile synthesis, and multimodal tactile frame prediction. We further propose \textbf{HandTouch}, a vector-quantized vision--tactile encoder that learns tactile representations through progressive spatial, cross-modal, and temporal training. Across HT-Bench, HandTouch consistently outperforms representative tactile encoder baselines, improving Recall@5 on fine-grained tactile similarity retrieval from 74.65\% to 85.23\%, reducing RMSE on masked tactile inpainting from 0.022 to 0.010, and increasing OOD cIoU on vision-to-tactile synthesis from 0.628 to 0.705. These results demonstrate the effectiveness of HandTouch and suggest that large-scale egocentric full-hand tactile data provides a scalable basis for evaluating and advancing tactile representation learning in dexterous manipulation.

cs.RO

Nanospheres with Patches Arranged in Polyhedrons from Self-Assembly of Solution-State Diblock Copolymers under Spherical Confinement

Self-assembly of sphere-forming solution-state amphiphilic diblock copolymers under spherical nanopore confinement is investigated using a simulated annealing technique. For two types of cases of different pore-surface/copolymer interactions, sequences of self-assembled patchy nanospheres are obtained, and phase diagrams are constructed. Self-assembled patchy nanospheres with 1-21 solvophobic domains are observed. The outermost solvophobic domains (patches) are packed into various polyhedrons when their number is larger than 3, where three Platonic solids of a regular tetrahedron, an octahedron, and an icosahedron and seven Johnson solids of J12, J13, J17, J50, J51, J86, and J87 are identified. In addition, another Johnson solid of J84 is identified in a structure with two categories of B-domains. These polyhedrons have all or most of their faces in a triangular shape, and hence, they are closer to spherical in shape, which may relieve the chain stretching. Nanospheres with 1, 4, 6, 9, and 12 numbers of patches occur in relatively large windows in the phase diagrams of both types of cases. In one of the two types of cases, all nanospheres with any number of 1-14 patches occur in the phase diagram, whereas in the other type of cases, nanospheres with 2, 3, 5, 11, and 13 numbers of patches are absent in the phase diagram. Furthermore, at a given pore size, the number of patches changes nonmonotonically or is unchanged with an increase in the strength of the pore-surface/copolymer interactions for one type or the other type of case, respectively. Quantitative calculations are performed to elucidate mechanisms of the window size in the phase diagrams of nanospheres with different numbers of patches and structure details.

cond-mat.soft