SearcharxivSearch

arXiv subjects

Jiasen Wang

Publications and source records attributed to Jiasen Wang.

5 recordsLinked to original sources

SuperSuit: An Isomorphic Bimodal Interface for Scalable Mobile Manipulation

High-quality, long-horizon demonstrations are essential for embodied AI, yet acquiring such data for tightly coupled wheeled mobile manipulators remains a fundamental bottleneck. Unlike fixed-base systems, mobile manipulators require continuous coordination between $SE(2)$ locomotion and precise manipulation, exposing limitations in existing teleoperation and wearable interfaces. We present \textbf{SuperSuit}, a bimodal data acquisition framework that supports both robot-in-the-loop teleoperation and active demonstration under a shared kinematic interface. Both modalities produce structurally identical joint-space trajectories, enabling direct data mixing without modifying downstream policies. For locomotion, SuperSuit maps natural human stepping to continuous planar base velocities, eliminating discrete command switches. For manipulation, it employs a strictly isomorphic wearable arm in both modes, while policy training is formulated in a shift-invariant delta-joint representation to mitigate calibration offsets and structural compliance without inverse kinematics. Real-world experiments on long-horizon mobile manipulation tasks show 2.6$\times$ higher demonstration throughput in active mode compared to a teleoperation baseline, comparable policy performance when substituting teleoperation data with active demonstrations at fixed dataset size, and monotonic performance improvement as active data volume increases. These results indicate that consistent kinematic representations across collection modalities enable scalable data acquisition for long-horizon mobile manipulation.

cs.RO

StreamVLA: Breaking the Reason-Act Cycle via Completion-State Gating

Long-horizon robotic manipulation requires bridging the gap between high-level planning (System 2) and low-level control (System 1). Current Vision-Language-Action (VLA) models often entangle these processes, performing redundant multimodal reasoning at every timestep, which leads to high latency and goal instability. To address this, we present StreamVLA, a dual-system architecture that unifies textual task decomposition, visual goal imagination, and continuous action generation within a single parameter-efficient backbone. We introduce a "Lock-and-Gated" mechanism to intelligently modulate computation: only when a sub-task transition is detected, the model triggers slow thinking to generate a textual instruction and imagines the specific visual completion state, rather than generic future frames. Crucially, this completion state serves as a time-invariant goal anchor, making the policy robust to execution speed variations. During steady execution, these high-level intents are locked to condition a Flow Matching action head, allowing the model to bypass expensive autoregressive decoding for 72% of timesteps. This hierarchical abstraction ensures sub-goal focus while significantly reducing inference latency. Extensive evaluations demonstrate that StreamVLA achieves state-of-the-art performance, with a 98.5% success rate on the LIBERO benchmark and robust recovery in real-world interference scenarios, achieving a 48% reduction in latency compared to full-reasoning baselines.

cs.RO

DVPE: Divided View Position Embedding for Multi-View 3D Object Detection

Sparse query-based paradigms have achieved significant success in multi-view 3D detection for autonomous vehicles. Current research faces challenges in balancing between enlarging receptive fields and reducing interference when aggregating multi-view features. Moreover, different poses of cameras present challenges in training global attention models. To address these problems, this paper proposes a divided view method, in which features are modeled globally via the visibility crossattention mechanism, but interact only with partial features in a divided local virtual space. This effectively reduces interference from other irrelevant features and alleviates the training difficulties of the transformer by decoupling the position embedding from camera poses. Additionally, 2D historical RoI features are incorporated into the object-centric temporal modeling to utilize highlevel visual semantic information. The model is trained using a one-to-many assignment strategy to facilitate stability. Our framework, named DVPE, achieves state-of-the-art performance (57.2% mAP and 64.5% NDS) on the nuScenes test set. Codes will be available at https://github.com/dop0/DVPE.

cs.CV

Flexible Design on Deterministic IP Networking for Mixed Traffic Transmission

Deterministic IP (DIP) networking is a promising technique that can provide delay-bounded transmission in large-scale networks. Nevertheless, DIP faces several challenges in the mixed traffic scenarios, including (i) the capability of ultra-low latency communications, (ii) the simultaneous satisfaction of diverse QoS requirements, and (iii) the network efficiency. The problems are more formidable in the dynamic surroundings without prior knowledge of traffic demands. To address the above-mentioned issues, this paper designs a flexible DIP (FDIP) network. In the proposed network, we classify the queues at the output port into multiple groups. Each group operates with different cycle lengths. FDIP can assign the time-sensitive flows with different groups, hence delivering diverse QoS requirements, simultaneously. The ultra-low latency communication can be achieved by specific groups with short cycle lengths. Moreover, the flexible scheduling with diverse cycle lengths improves resource utilization, hence increasing the throughput (i.e., the number of acceptable time-sensitive flows). We formulate a throughput maximization problem that jointly considers the admission control, transmission path selection, and cycle length assignment. A branch and bound (BnB)-based heuristic is developed. Simulation results show that the proposed FDIP significantly outperforms the standard DIP in terms of both the throughput and the latency guarantees.

cs.NI

Achieving Deterministic Service in Mobile Edge Computing (MEC) Networks

Mobile edge computing (MEC) is proposed to boost high-efficient and time-sensitive 5G applications. However, the "microburst" may occur even in lightly-loaded scenarios, which leads to the indeterministic service latency (i.e., unpredictable delay or delay variation), hence hindering the deployment of MEC. Deterministic IP networking (DIP) has been proposed that can provide bounds on latency, and high reliability in the large-scale networks. Nevertheless, the direct migration of DIP into the MEC network is non-trivial owing to its original design for the Ethernet with homogeneous devices. Meanwhile, DIP also faces the challenges on the network throughput and scheduling flexibility. In this paper, we delve into the adoption of DIP for the MEC networks and some of the relevant aspects. A deterministic MEC (D-MEC) network is proposed to deliver the deterministic service (i.e., providing the MEC service with bounded service latency). In the D-MEC network, two mechanisms, including the cycle mapping and cycle shifting, are designed to enable: (i) seamless and deterministic transmission with heterogeneous underlaid resources; and (ii) traffic shaping on the edges to improve the resource utilization. We also formulate a joint configuration to maximize the network throughput with deterministic QoS guarantees. Extensive simulations verify that the proposed D-MEC network can achieve a deterministic MEC service, even in the highly-loaded scenarios.

cs.NI