arXiv · 2605.25806
An Analysis Focused on Womens Safety: Can VAD Models Be Enhanced by a Multi-modal Dataset?
Abstract
Women's safety and security are paramount for a modern society. Often, crimes scenes get recorded through low-resolution CCTV cameras limiting the efficiency of video anomaly detection (VAD) models. Despite substantial progress in VAD research, women-centric anomalies are still underrepresented in datasets as well as in models. Existing datasets primarily cover well-lit, high-resolution and close-shot videos that are inadequate to tackle critical anomalies such as chain snatching, stalking, inappropriate touch, and other subtle forms of crime against women. To address this, we present a new benchmark, referred to as ExtrAnom. It contains 1001 videos (both anomalies and normal) with four textual annotations; one human-generated and three LLM-generated. The videos are arranged in 5 different categories of crimes. The dataset comprises low-light (8%), low-resolution (13%), long-shot (15%), and daytime (64%) anomaly videos. It includes stalking (3.9%), chain snatching (17.6%), kidnapping (7.3%), assassinations (2.3%), harassment (18.9%), and normal (50%) videos. It is possible to perform cross-modal and VLM-based validations using ExtrAnom. We have benchmarked it against popular unimodal and multi-modal VAD datasets (e.g., XD-Violence, UCF-Crime, and UCA) and SOTA methods. Experiments reveal that existing datasets are insufficient to deal with women-centric anomalies. We believe ExtrAnom can fill this critical gap in VAD research.
Explore related subjects
Keep this discovery
Sangeeta ., Maddikuntla Sai Prajwal, Debi Prosad Dogra, Kamalakar Vijay Thakare, Hyungjoo Jung, Ig-Jae Kim, Heeseung Choi. 2026-05-25. An Analysis Focused on Womens Safety: Can VAD Models Be Enhanced by a Multi-modal Dataset?. https://arxiv.org/abs/2605.25806
Cite the original work for its findings. Save a collection to share your selection of sources.