SearcharxivSearch

arXiv subjects

Kirato Yoshihara

Publications and source records attributed to Kirato Yoshihara.

2 recordsLinked to original sources

VesselBridge3D: A Foundation Model Adaptation Framework for Label-Efficient 3D Vessel Segmentation

State-of-the-art vessel segmentation methods typically require large-scale annotated datasets and suffer from severe performance degradation under domain shifts. In clinical practice, however, acquiring extensive annotations for every new scanner or protocol is unfeasible. To address this, we propose VesselBridge3D, a foundation model adaptation framework that bridges frozen vision foundation models and volumetric vessel segmentation through lightweight 3D adaptation modules. The framework combines a lightweight 3D Adapter, a multi-scale 3D Aggregator, and Z-channel embedding for efficient adaptation to volumetric medical images. We instantiate VesselBridge3D with three frozen foundation encoders (DINOv3, MedSAM, and MedGemma) and evaluate it on the TopCoW (ID) and Lausanne (OOD) datasets. In the extreme low-data regime with 5 training samples, our method achieved a Dice score of 43.42%, marking a 30% relative improvement over the state-of-the-art nnU-Net (33.41%) and outperforming other Transformer-based baselines by up to 45%. The proposed framework was effective across all evaluated frozen foundation encoders, with DINOv3 yielding the best performance in the most label-efficient settings. Furthermore, in the out-of-distribution setting, our model demonstrated superior robustness, achieving a 50% relative improvement over nnU-Net (21.37% vs. 14.22%), which suffered from severe domain overfitting. Ablation studies confirmed the effectiveness of the proposed 3D adaptation modules. Our results demonstrate that VesselBridge3D is an effective framework for label-efficient 3D vessel segmentation under data scarcity and domain shifts.

eess.IV

Different Layers, Different Manifolds: Module-Wise Weight-Space Geometry in Transformer Optimization

Weight-space geometry plays a central role in neural network optimization, yet manifold constraints are often applied uniformly across all weight matrices. In this work, we ask whether different transformer modules prefer different manifold geometries. We study Manifold Muon for GPT-2 pretraining and compare layer-wise assignments of Stiefel and DGram constraints across attention and MLP blocks. Our results show a clear asymmetry: constraining attention layers with Stiefel geometry while assigning DGram geometry to MLP layers gives the best performance among the tested configurations, whereas the inverted assignment and all-DGram configuration become unstable under the shared hyperparameter setting. We trace this failure to singular value growth in DGram-constrained attention weights, which can amplify attention logits and induce softmax saturation. These findings suggest that symmetry-aware and geometry-aware optimization for transformers should be module-specific rather than uniform.

cs.LG