arXiv · 2407.06178
Transfer Learning with Self-Supervised Vision Transformers for Snake Identification
Abstract
We present our approach for the SnakeCLEF 2024 competition to predict snake species from images. We explore and use Meta's DINOv2 vision transformer model for feature extraction to tackle species' high variability and visual similarity in a dataset of 182,261 images. We perform exploratory analysis on embeddings to understand their structure, and train a linear classifier on the embeddings to predict species. Despite achieving a score of 39.69, our results show promise for DINOv2 embeddings in snake identification. All code for this project is available at https://github.com/dsgt-kaggle-clef/snakeclef-2024.
Explore related subjects
Keep this discovery
Anthony Miyaguchi, Murilo Gustineli, Austin Fischer, Ryan Lundqvist. 2024-07-08. Transfer Learning with Self-Supervised Vision Transformers for Snake Identification. https://arxiv.org/abs/2407.06178
Cite the original work for its findings. Save a collection to share your selection of sources.