arXiv · 2607.20291
Diverse-Intent Multi-Turn Fashion Image Retrieval
Abstract
Real-world fashion search involves interactive retrieval across multiple turns. However, existing multi-turn retrieval methods are built on a restrictive assumption that every interaction follows the same attribute-editing paradigm, leaving heterogeneous intent transitions unexplored. Moreover, existing approaches often rely on textification to bridge multimodal queries and visual retrieval, which may lose fine-grained visual cues. To address these gaps, we introduce DIM-Fashion, a benchmark of 26K multi-turn sessions constructed from 13 fashion retrieval datasets across 7 tasks, featuring diverse intent transitions and rollback behaviors. We further propose FashionAM, an MLLM-VLP framework that directly aligns multimodal conversational queries with a fashion-oriented gallery embedding space, avoiding intermediate textification. Extensive experiments demonstrate the effectiveness of FashionAM over existing approaches. The dataset and code will be made publicly available upon acceptance.
Explore related subjects
Keep this discovery
Mingqiang Tang, Haokun Wen, Meng Liu, Yupeng Hu, Weili Guan, Xuemeng Song. 2026-07-22. Diverse-Intent Multi-Turn Fashion Image Retrieval. https://arxiv.org/abs/2607.20291
Cite the original work for its findings. Save a collection to share your selection of sources.