arXiv · 2504.20322
Fine-Grained Classification: Connecting Metadata via Cross-Contrastive Pre-Training
Abstract
Fine-grained visual classification aims to recognize objects belonging to many subordinate categories of a supercategory, where appearance alone often fails to distinguish highly similar classes. We propose a unified framework that integrates image, text, and metadata via cross-contrastive pre-training. We first align the three modality encoders in a shared embedding space and then fine-tune the image and metadata encoders for classification. On NABirds, our approach improves over the baseline by 7.83% and achieves 84.44% top-1 accuracy, outperforming strong multimodal methods.
Explore related subjects
Keep this discovery
Sumit Mamtani, Yash Thesia. 2025-04-29. Fine-Grained Classification: Connecting Metadata via Cross-Contrastive Pre-Training. https://arxiv.org/abs/2504.20322
Cite the original work for its findings. Save a collection to share your selection of sources.