Searcharxiv⌕ Search

arXiv · 2609.22726

Decentralized Multi-Robot Exploration with Probabilistic Peer Intent and Multi-hop Plan Propagation

Abstract

Efficient coordination under limited communication remains a key challenge in decentralized multi-robot exploration. While centralized approaches benefit from global information sharing, they are often impractical in large-scale or communication-constrained environments. Existing Monte Carlo Tree Search (MCTS)-based approaches, such as Decentralized Monte Carlo Exploration (DMCE), enable decentralized planning by taking peer intent into account. This peer intent is obtained by communicating sequences of planned waypoints with robots within direct communication range. In this work, we extend this idea by introducing Probabilistic Peer Intent (PPI), which converts peer trajectories into a continuous spatial representation of predicted intent and incorporates it into local MCTS action evaluation. We additionally study the effects of sharing peer intent beyond direct communication range by propagating plans over multiple hops. Experiments across multiple simulated environments and team sizes show that PPI and Multi-hop propagation can each improve decentralized exploration, with their relative benefits depending on environment structure and team size. We also demonstrate the real-world deployment of our method on three robots operating in different environment types.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Saurbh Singh Jamwal, Nived Chebrolu, Shivaram Kalyanakrishnan. 2026-09-19. Decentralized Multi-Robot Exploration with Probabilistic Peer Intent and Multi-hop Plan Propagation. https://arxiv.org/abs/2609.22726

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Easier Said Than Done: Unpacking Intent-Behavior Gap in Jailbreaking LLM-based Robots

LLM-based robots use Large Language Models (LLMs) as planners to translate natural language instructions into policies such as grasp(), move_to(), and open_gripper(). Jailbreak attacks on these robots extend the threat from generating malicious content to executing harmful behaviors. However, we find that existing jailbreak attempts against LLM-based robots that produce malicious-looking policies (intent jailbreaks) often fail to induce harmful physical actions by robots (behavior jailbreaks), due to robot-specific constraints, such as logical errors and hallucinated control APIs. In this paper, we demystify the intent-behavior gap and investigate its root causes to inform effective defenses. Our measurement study finds that current LLM jailbreak methods overlook robot-specific syntax constraints (e.g., executable control APIs) and physical feasibility (e.g., ordering of policies and hardware/kinematic constraints). To bridge the gap, we introduce POEF (Policy Effective Jailbreak), an automated red-teaming framework that takes into account the robot-specific constraints during both the optimization and evaluation processes. Specifically, POEF employs the hidden-layer gradients from an unaligned LLM to guide the jailbreak prompt optimization and uses a multi-agent evaluator to assess the feasibility of the generated policies. Experiments on commercial robots, including the Unitree G1, the Franka robotic arm, and simulators, show that POEF achieves an 80% behavior jailbreak success rate and transfers across various LLMs. In addition, we propose two defense strategies that mitigate the behavior jailbreak risks. Our findings indicate an urgent need for stronger countermeasures before LLM-based robots are deployed at scale. The homepage is available at https://zjushine.github.io/poef.github.io/.

cs.RO↗

SPIDER: Scalable Physics-Informed Dexterous Retargeting

Learning agile robotic policies requires large-scale demonstrations, but translating abundant human motion data to robots is bottlenecked by the embodiment gap and missing dynamic information. To bridge this gap, we propose Scalable Physics-Informed DExterous Retargeting (SPIDER), a physics-based retargeting framework to transform and augment kinematic-only human demonstrations into dynamically feasible robot trajectories at scale. Our key insight is that human demonstrations should provide global task structure and objective, while a sampling-based solver can be used to find the feasible solution given the physics constraints. As a general framework, SPIDER is an efficient physics-based retargeting method that can be applied to both humanoid whole-body loco-manipulation and dexterous manipulation across 9 humanoid/dexterous hand embodiments, 6 datasets and 3 simulators. By bypassing the need for policy optimization, it achieves state-of-the-art performance while being 10x faster than reinforcement learning (RL) baselines. Furthermore, SPIDER enables physics-based data augmentation, such as imposing external payloads, perturbations, and new contact patterns, to generate diverse data. We demonstrate that the retargeted motion can be executed on the robot directly open-loop or serve as feasible reference motion for efficient RL policy learning. Our pipeline enables large-scale generation of robot trajectories from human motion, and we release a dataset of 6,520 simulated demonstrations across four hands and 191 dataset-specific object identities to support future research.

cs.RO↗

Learning Control Policies to Provably Satisfy Hard Affine Constraints for Black-Box Hybrid Dynamical Systems

Ensuring safety for black-box hybrid dynamical systems presents significant challenges due to their instantaneous state jumps and unknown explicit nonlinear dynamics. Existing solutions for strict safety constraint satisfaction, like control barrier functions (CBFs) and reachability analysis, rely on direct knowledge of the dynamics. Similarly, safe reinforcement learning (RL) approaches often rely on known system dynamics or merely discourage safety violations through reward shaping. In this work, we want to learn RL policies which provably satisfy affine state constraints in closed loop for black-box hybrid dynamical systems with affine reset maps. Our key insight is forcing the RL policy to be affine and repulsive near the constraint boundaries for the unknown nonlinear dynamics of the system, providing guarantees that the trajectories will not violate the constraint. We further account for constraint violation due to instantaneous state jumps that occur due to impacts or reset maps in the hybrid system by introducing a second repulsive affine region before the reset that prevents post-reset states from violating the constraint. We derive sufficient conditions under which these policies satisfy safety constraints in closed loop. We also compare our approach with state-of-the-art reward shaping and learned-CBF methods on hybrid dynamical systems like the constrained pendulum and paddle juggler environments. In both scenarios, we show that our methodology learns higher quality policies while always satisfying the safety constraints.

cs.RO↗