arXiv · 2601.22376
FlexMap: Robust HD Map Construction under Flexible Camera Configurations
Abstract
High-definition (HD) maps provide essential semantic information about road structures for autonomous driving, but existing HD map construction methods typically require calibrated multi-camera rigs and explicit 2D-to-BEV transformations. Such pipelines degrade when camera views are missing or pose estimates are inaccurate, limiting their use across heterogeneous fleet configurations. We introduce FlexMap, a vectorized HD mapping framework that adapts to varying camera configurations without architectural changes or per-configuration retraining and does not require camera parameters as model input. FlexMap replaces explicit geometric projection with a geometry foundation model that encodes cross-view 3D structure. A spatial-temporal enhancement module then separates cross-view spatial reasoning from temporal aggregation, while a camera-aware decoder uses the token produced for each input view to adapt its attention without camera poses. Experiments on nuScenes and Argoverse 2 show that FlexMap outperforms pose-dependent baselines using estimated poses and maintains comparable accuracy across all evaluated camera configurations, including those with missing views.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Run Wang, Chaoyi Zhou, Amir Salarpour, Xi Liu, Zhi-Qi Cheng, Feng Luo, Mert D. Pesé, Siyu Huang. 2026-01-29. FlexMap: Robust HD Map Construction under Flexible Camera Configurations. https://arxiv.org/abs/2601.22376
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.