arXiv · 2608.21244
A VLM Answer Is Not an Anomaly Score: Rank Compression Across Image and Video Anomaly Detection
Abstract
Anomaly detection aims to identify observations that deviate from normal patterns. Recent work uses pretrained vision-language models (VLMs) for training-free image and video anomaly detection without task-specific retraining. Anomaly detection is commonly evaluated by how well anomaly scores rank anomalous images or video frames above normal ones. Generative VLMs, however, assign probabilities to possible answers and then decode a single answer. This decoding step can discard ordering information. We call this loss of ordering decoded-answer rank compression and study whether it materially affects anomaly detection performance. To isolate this effect, we compare two ways of scoring the same VLM output: one uses only the decoded answer, while the other computes a probability-weighted score over all possible answers. Across image and video anomaly detection benchmarks, VLMs, and answer scales, probability-weighted scoring consistently outperforms decoded-answer scoring, with mean gains ranging from 7.66 to 19.95 points on the primary benchmark metrics. Using answer probabilities only to break ties created by decoded-answer scoring recovers at least 95% of the average performance gap on every benchmark. When answer probabilities are available, how VLM answers are converted into anomaly scores is therefore part of the detector design, not merely an implementation detail.
Explore related subjects
Keep this discovery
Inpyo Song, Jangwon Lee. 2026-08-21. A VLM Answer Is Not an Anomaly Score: Rank Compression Across Image and Video Anomaly Detection. https://arxiv.org/abs/2608.21244
Cite the original work for its findings. Save a collection to share your selection of sources.