arXiv · 2509.08374
InsFusion: Rethink Instance-level LiDAR-Camera Fusion for 3D Object Detection
Abstract
Three-dimensional Object Detection from multi-view cameras and LiDAR is a crucial component for autonomous driving and smart transportation. However, in the process of basic feature extraction, perspective transformation, and feature fusion, noise and error will gradually accumulate. To address this issue, we propose InsFusion, which can extract proposals from both raw and fused features and utilizes these proposals to query the raw features, thereby mitigating the impact of accumulated errors. Additionally, by incorporating attention mechanisms applied to the raw features, it thereby mitigates the impact of accumulated errors. Experiments on the nuScenes dataset demonstrate that InsFusion is compatible with various advanced baseline methods and delivers new state-of-the-art performance for 3D object detection.
Explore related subjects
Keep this discovery
Zhongyu Xia, Hansong Yang, Yongtao Wang. 2025-09-10. InsFusion: Rethink Instance-level LiDAR-Camera Fusion for 3D Object Detection. https://arxiv.org/abs/2509.08374
Cite the original work for its findings. Save a collection to share your selection of sources.