arXiv · 2511.11427
Comprehension of Multilingual Expressions Referring to Target Objects in Visual Inputs
Abstract
Referring Expression Comprehension (REC) requires models to localize objects in images based on different types of natural language descriptions. Even with significant progress, research on the area remains predominantly English-centric, despite increasing global deployment demands. This work addresses multilingual REC through two main contributions. First, we construct a unified multilingual dataset spanning 10 languages, by systematically expanding 12 existing English REC benchmarks through machine translation and context-based translation enhancement. Second, we introduce an attention-anchored efficient neural architecture that uses a multilingual SigLIP2 encoder. Our attention-based approach generates coarse spatial anchors from attention distributions, which are subsequently refined through learned residuals. Experimental evaluation demonstrates competitive performance on standard benchmarks despite the use of a relatively small model. Multilingual evaluation shows consistent capabilities across languages, establishing the practical feasibility of efficient multilingual visual grounding systems.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Francisco Nogueira, Alexandre Bernardino, Bruno Martins. 2025-11-14. Comprehension of Multilingual Expressions Referring to Target Objects in Visual Inputs. https://arxiv.org/abs/2511.11427
Cite the original work for its findings. Save a collection to share your selection of sources.