arXiv · 2501.05488
EndoDINO: A Foundation Model for GI Endoscopy
Abstract
In this work, we present EndoDINO, a foundation model for GI endoscopy tasks that achieves strong generalizability by pre-training on a well-curated image dataset sampled from the largest known GI endoscopy video dataset in the literature. Specifically, we pre-trained ViT models with 1B, 307M, and 86M parameters using datasets ranging from 100K to 10M curated images. Using EndoDINO as a frozen feature encoder, we achieved state-of-the-art performance in anatomical landmark classification, polyp segmentation, and Mayo endoscopic scoring (MES) for ulcerative colitis with only simple decoder heads.
Explore related subjects
Keep this discovery
Patrick Dermyer, Angad Kalra, Matt Schwartz. 2025-01-08. EndoDINO: A Foundation Model for GI Endoscopy. https://arxiv.org/abs/2501.05488
Cite the original work for its findings. Save a collection to share your selection of sources.