arXiv · 2610.08407
Seeing the Context: Enhancing Recommender Systems with Image-Derived Contextual Signals
Abstract
Contextual information, capturing the circumstances of a user-item interaction, is central to recommender systems. Prior work draws context from location, time, or reviews, but not images; multimodal recommender systems mainly use images to enrich item or user representations, not identify situational context. We propose a new representation of context derived from images, spanning physical, social, and modal categories learned via a vision-language model. We introduce ICE-Fuse, a pipeline for evaluating this representation that fuses these categories and integrates them into a context-aware recommender system, using TripAdvisor data and Review-aware Graph Contrastive Learning as the recommendation algorithm. Image context does not outperform established signals standalone, but improves them combined, indicating complementary information. Semantic analysis shows image- and review-derived context capture distinct aspects of the interaction, positioning images as complementary context.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Tal Cordova, Tomer Geva, Moshe Unger. 2026-10-06. Seeing the Context: Enhancing Recommender Systems with Image-Derived Contextual Signals. https://arxiv.org/abs/2610.08407
Cite the original work for its findings. Save a collection to share your selection of sources.