arXiv · 2609.34781
When VLMs Trust Context: Evaluating Scene Text Recognition under Misleading Context
Abstract
Vision-language models (VLMs) can read text in natural scenes, but their predictions may be influenced by the surrounding context. When the printed text conflicts with what the scene suggests, a model may return a more plausible word instead of the shown text. We introduce SceneFaith, a benchmark of 781 generated scene images for studying this behavior. Each output is classified as Literal, Canonical, or Other, separating faithful transcription from context-consistent rewriting and ordinary recognition errors. Across 15 models from seven families, all models show rewriting on clear images, with rates ranging from 8.45\% to 58.51\%. Controlled experiments further show that surrounding context matters: removing surrounding scene information reduces rewriting and improves literal accuracy, while changing the scene around the same text patch can also change model outputs. Moreover, weakening the target text with blur increases rewriting. These results show that reliable scene-text recognition requires VLMs to balance visual character evidence with contextual information, preserving clear text while using context mainly when the visual evidence is uncertain.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yuxing Cheng, Yuan Wu, Yi Chang. 2026-09-28. When VLMs Trust Context: Evaluating Scene Text Recognition under Misleading Context. https://arxiv.org/abs/2609.34781
Cite the original work for its findings. Save a collection to share your selection of sources.