arXiv · 2609.32876
Multimodal LLMs Outperform Pathology Foundation Models in Cross-Domain Histological Similarity
Abstract
State-of-the-art pathology foundation models, trained on millions of histology tiles, can fail to preserve tissue similarity when comparisons cross slide or institution boundaries. We show that general-purpose multimodal LLMs, without being trained as pathology foundation models, consistently outperform these specialized models in cross-domain histological similarity judgments. Using a relative similarity framework that we release as the MOSAIC (Model Similarity Assessment across Institutions and Cohorts) benchmark, we evaluate 17 models across 6 datasets and find that pathology encoders often rank same-institution, different-disease tiles as more similar than same-disease, different-institution tiles, a clinically dangerous failure mode invisible to standard within-domain evaluations. LLMs appear less susceptible to this failure, likely because they perform semantic visual comparison of morphology and tissue architecture rather than relying on shortcut features tied to acquisition context. Scaling training data does not resolve the problem for pathology encoders, implicating the learning objective rather than data coverage. Our results expose a fundamental robustness gap in current pathology foundation models and establish multimodal LLMs as a viable alternative for cross-institutional retrieval, dataset harmonization, and multi-site quality control. Code and data will be released upon acceptance.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yishu Zhang, Yun Li, Daiwei Zhang. 2026-09-26. Multimodal LLMs Outperform Pathology Foundation Models in Cross-Domain Histological Similarity. https://arxiv.org/abs/2609.32876
Cite the original work for its findings. Save a collection to share your selection of sources.