SearcharxivSearch

arXiv subjects

Yibo Ai

Publications and source records attributed to Yibo Ai.

2 recordsLinked to original sources

End-to-End 3-D Spatiotemporal Perception with Multimodal Fusion and V2X Collaboration

Multiview cooperative perception and multimodal fusion are essential for reliable 3-D spatiotemporal understanding in autonomous driving, especially in cases with occlusions, limited viewpoints, and communication delays in vehicle-to-everything (V2X) scenarios. In this paper, Cross-modal End-to-End Tracking for V2X (XET-V2X), a multimodal fused end-to-end tracking framework for V2X collaboration that unifies multiview multimodal sensing within a shared spatiotemporal representation, is proposed. To efficiently align heterogeneous viewpoints and modalities, XET-V2X introduces a dual-layer spatial cross-attention module based on multiscale deformable attention. Multiview image features are aggregated to enhance semantic consistency, followed by point cloud fusion guided by the updated spatial queries, enabling effective cross-modal interaction while reducing computational overhead. Experiments based on the real-world V2X Sequential Perception Dataset (V2X-Seq-SPD) dataset and two simulated V2X-Sim-derived subsets, namely the vehicle-to-vehicle (V2X-Sim-V2V) and vehicle-to-infrastructure (V2X-Sim-V2I) subsets, demonstrate consistent improvements in detection and tracking performance under varying communication delays, with XET-V2X achieving up to 15-20% relative gains in mean average precision (mAP) and average multi-object tracking accuracy (AMOTA) over single-view or single-modal baselines, while also outperforming representative tracking-by-detection cooperative perception methods.

cs.CV

LiDAR-based End-to-end Temporal Perception for Vehicle-Infrastructure Cooperation

Temporal perception, defined as the capability to detect and track objects across temporal sequences, serves as a fundamental component in autonomous driving systems. While single-vehicle perception systems encounter limitations, stemming from incomplete perception due to object occlusion and inherent blind spots, cooperative perception systems present their own challenges in terms of sensor calibration precision and positioning accuracy. To address these issues, we introduce LET-VIC, a LiDAR-based End-to-End Tracking framework for Vehicle-Infrastructure Cooperation (VIC). First, we employ Temporal Self-Attention and VIC Cross-Attention modules to effectively integrate temporal and spatial information from both vehicle and infrastructure perspectives. Then, we develop a novel Calibration Error Compensation (CEC) module to mitigate sensor misalignment issues and facilitate accurate feature alignment. Experiments on the V2X-Seq-SPD dataset demonstrate that LET-VIC significantly outperforms baseline models. Compared to LET-V, LET-VIC achieves +15.0% improvement in mAP and a +17.3% improvement in AMOTA. Furthermore, LET-VIC surpasses representative Tracking by Detection models, including V2VNet, FFNet, and PointPillars, with at least a +13.7% improvement in mAP and a +13.1% improvement in AMOTA without considering communication delays, showcasing its robust detection and tracking performance. The experiments demonstrate that the integration of multi-view perspectives, temporal sequences, or CEC in end-to-end training significantly improves both detection and tracking performance. All code will be open-sourced.

cs.CV