arXiv · 2306.02928
LRVS-Fashion: Extending Visual Search with Referring Instructions
Abstract
This paper introduces a new challenge for image similarity search in the context of fashion, addressing the inherent ambiguity in this domain stemming from complex images. We present Referred Visual Search (RVS), a task allowing users to define more precisely the desired similarity, following recent interest in the industry. We release a new large public dataset, LRVS-Fashion, consisting of 272k fashion products with 842k images extracted from fashion catalogs, designed explicitly for this task. However, unlike traditional visual search methods in the industry, we demonstrate that superior performance can be achieved by bypassing explicit object detection and adopting weakly-supervised conditional contrastive learning on image tuples. Our method is lightweight and demonstrates robustness, reaching Recall at one superior to strong detection-based baselines against 2M distractors. The dataset is available at https://huggingface.co/datasets/Slep/LAION-RVS-Fashion .
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Simon Lepage, Jérémie Mary, David Picard. 2023-06-05. LRVS-Fashion: Extending Visual Search with Referring Instructions. https://arxiv.org/abs/2306.02928
Cite the original work for its findings. Save a collection to share your selection of sources.