AD-SAM: Adapting the Segment Anything Model for Semantic Segmentation in Autonomous Driving
This paper presents the Autonomous Driving Segment Anything Model (AD-SAM), a foundation-model adaptation framework for semantic segmentation in autonomous driving. AD-SAM combines a frozen Segment Anything Model (SAM) Vision Transformer (ViT-H) encoder with a trainable ResNet-50 encoder to integrate general visual representations with multi-scale, domain-specific spatial features. Features from the two encoders are integrated through deformable convolution and channel attention, followed by a multi-stage deformable decoder for semantic prediction. Training employs a hybrid objective combining Focal, Dice, Lovász-Softmax, and Surface losses. Experiments on Cityscapes and Berkeley DeepDrive 100K (BDD100K) show that AD-SAM outperforms SAM, Generalized SAM (G-SAM), and DeepLabV3 under a controlled training protocol. AD-SAM achieves 76.27% mIoU on Cityscapes and 64.74% on BDD100K, exceeding DeepLabV3 by 3.45 and 5.00 percentage points, respectively, with larger gains over the SAM-based baselines. Sample-size experiments reveal dataset-dependent behavior. In particular, AD-SAM performs strongly across training sizes on Cityscapes, while its advantage on the more heterogeneous BDD100K becomes pronounced with increased training data. When trained on Cityscapes and directly evaluated on BDD100K, AD-SAM achieves the highest cross-dataset retention (84.88%) among the evaluated models. AD-SAM also converges rapidly, while precomputed frozen SAM embeddings reduce training memory requirements. These findings demonstrate the potential of combining general foundation-model representations with domain-specific multi-scale features for accurate and robust autonomous-driving semantic segmentation.