Searcharxiv⌕ Search

arXiv · 2610.10283

Temporal Visuo-Tactile Learning for Dexterous Grasp Stability

Abstract

Humans can grasp everyday objects with almost perfect success rates using fingertip tactile feedback, yet much of the robotic grasping literature emphasizes vision-based grasp selection with parallel grippers. In this work, we systematically investigate how high-resolution, dynamic tactile sensing contributes to grasp stability prediction and model-guided grasping in dexterous robotic hands. To this end, we collected a dataset of 10,000 grasp trials across 200 objects using a multi-fingered robotic hand equipped with four Digit 360 tactile sensors, recording external vision, proprioception, and tactile streams throughout each grasp. With this dataset, we trained end-to-end temporal multimodal models to predict post-lift stability from pre-lift grasp observations and compared sensing modalities and encoding backbones. Experimental results and controlled input ablations show that incorporating touch, and particularly high-resolution, dynamic touch, improves grasp stability prediction. Finally, we deployed the learned predictor as an online stability gate on the real robot, where visuo-tactile model-guided regrasping improved the success rate among executed lifts by 10.5 percentage points over a non-tactile gate. These results show how rich fingertip sensing and expressive temporal models that capture the dynamics of touch can support learned grasping with multi-fingered hands without explicit contact or force modeling, providing a scalable data-driven path from tactile experience toward stable dexterous manipulation. The dataset is publicly available at https://lasr-lab.github.io/dexterous-grasp-stability/.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ken Nakahara, Aleksei Buvailik, Prokhor Kotov, Roberto Calandra. 2026-10-07. Temporal Visuo-Tactile Learning for Dexterous Grasp Stability. https://arxiv.org/abs/2610.10283

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

RoboAug: One Annotation to Hundreds of Scenes via Region-Contrastive Data Augmentation for Robotic Manipulation

Enhancing the generalization of robotic learning in diverse unseen environments remains a fundamental challenge. Existing approaches often rely on large-scale pretraining, which is labor-intensive and time-consuming, or semantic data augmentation methods that assume flawless upstream object detection in real-world scenarios. In this work, we propose RoboAug, a novel generative data augmentation framework that reduces reliance on large-scale pretraining and perfect visual recognition by requiring only a single image with bounding box annotations for dataset construction. Leveraging this minimal supervision, RoboAug employs pretrained generative models for precise semantic augmentation and introduces a plug-and-play region-contrastive loss to guide attention toward task-relevant regions, thereby enhancing generalization and task success rates. Extensive real-world experiments on UR-5e, AgileX, and Tian Gong 2.0 demonstrate that RoboAug consistently outperforms state-of-the-art augmentation baselines under background, distractor, and lighting shifts. Our project is available at https://x-roboaug.github.io/.

cs.RO↗

Seed2Scale: A Self-Evolving Data Engine with Parallel Worlds Expansion for Scalable Robot Learning

Existing data generation methods for robot learning suffer from limited exploration, embodiment gaps, low signal-to-noise ratios, and domain shifts, leading to performance degradation during self-iteration and poor generalization to unseen scenes. To address these challenges, we propose Seed2Scale, a self-evolving data engine with parallel worlds expansion. Starting with as few as four seed demonstrations, Seed2Scale first executes a self-evolution stage driven by a heterogeneous synergy of "small-model collection, large-model evaluation, and target-model learning". Specifically, the lightweight Vision-Language-Action (VLA) model, SuperTiny, serves as a dedicated data collector for robust exploration. Concurrently, a pretrained Vision-Language Model (VLM) functions as a verifier to autonomously score and filter trajectories, supporting stable self-evolution in the evaluated tasks without performance collapse. Furthermore, Seed2Scale introduces a parallel worlds stage, projecting self-evolved trajectories into different environments of the same task to generate more diverse data and enhance adaptability to unseen scenes, including real-world environments. Experimental results demonstrate that Seed2Scale exhibits significant scaling potential: as iterations progress, the success rate of the target model shows a consistent upward trend, significantly outperforming the seed baseline. Notably, Seed2Scale achieves a remarkable 75.38% success rate in zero-shot real-world evaluations, where baseline methods fail completely (0%). Project page: https://terminators2025.github.io/Seed2Scale.github.io

cs.RO↗

SPAN-Nav: Generalized Spatial Awareness for Versatile Embodied Navigation

Recent embodied navigation approaches leveraging Vision-Language Models (VLMs) demonstrate strong generalization in versatile Vision-Language Navigation (VLN). However, reliable path planning in complex environments remains challenging due to insufficient spatial awareness. In this work, we introduce SPAN-Nav, an end-to-end foundation model designed to infuse embodied navigation with universal 3D spatial awareness using RGB video streams. SPAN-Nav extracts spatial priors across diverse scenes through an occupancy prediction task on extensive indoor and outdoor environments. To mitigate the computational burden, we introduce a compact representation for spatial priors, finding that a single token is sufficient to encapsulate the coarse-grained cues essential for navigation tasks. Furthermore, inspired by the Chain-of-Thought (CoT) mechanism, SPAN-Nav utilizes this single spatial token to explicitly inject spatial cues into action reasoning through an end-to end framework. Leveraging multi-task co-training, SPAN-Nav captures task-adaptive cues from generalized spatial priors, enabling robust spatial awareness to generalize even to the task lacking explicit spatial supervision. To support comprehensive spatial learning, we present a massive dataset of 4.2 million occupancy annotations that covers both indoor and outdoor scenes across multi-type navigation tasks. SPAN-Nav achieves state-of-the-art performance across three benchmarks spanning diverse scenarios and varied navigation tasks. Finally, real-world experiments validate the robust generalization and practical reliability of our approach across complex physical scenarios.

cs.RO↗