arXiv · 2511.05782
TCSA-UDA: Text-Driven Cross-Semantic Alignment for Unsupervised Domain Adaptation in Medical Image Segmentation
Abstract
Unsupervised domain adaptation (UDA) for medical image segmentation remains challenging due to substantial domain shifts across imaging modalities, such as CT and MRI. Although recent vision-language representation learning methods have shown promise in medical image analysis, their role in cross-modality UDA segmentation remains underexplored. To address this problem, we propose TCSA-UDA, a Text-driven Cross-Semantic Alignment framework that uses modality-aware textual prompting to guide domain-invariant visual representation learning. Specifically, we introduce a vision-language covariance cosine loss (VLCoL) that aligns inter-class visual feature relationships with text-derived semantic relationships, encouraging the image encoder to learn semantically structured and modality-robust representations. In addition, we incorporate a prototype alignment module to reduce residual class-level discrepancies between source and target domains by aligning high-level class prototypes. Extensive experiments on cross-modality cardiac, abdominal, and brain tumor segmentation benchmarks demonstrate that TCSA-UDA consistently improves adaptation performance and outperforms state-of-the-art UDA methods. These results highlight the potential of language-driven semantic guidance for domain-adaptive medical image segmentation. The code is available at https://github.com/lalitmaurya47/TCSA_UDA
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Lalit Maurya, Honghai Liu, Reyer Zwiggelaar. 2025-11-08. TCSA-UDA: Text-Driven Cross-Semantic Alignment for Unsupervised Domain Adaptation in Medical Image Segmentation. https://doi.org/10.1016/j.neucom.2026.134808
Cite the original work for its findings. Save a collection to share your selection of sources.