Beyond Isolated Entities: Relation-Aware Multi-Entity Modeling for Unsupervised Video Anomaly Detection
Video anomaly detection for patrol robots and surveillance systems must recognize abnormal interactions among familiar entities. Existing pixel-reconstruction and isolated-entity methods may fail when individual entities appear normal but their spatial or motion relations are abnormal. This work presents Interaction-Centric Network for Temporal Entity-Relation Analysis and Consistency Testing (INTERACT), a framework for unsupervised video anomaly detection. Person-Object Appearance-Motion Interaction (POAMI) moves beyond isolated-entity encoding by jointly modeling the appearance, motion, and spatial configurations of persons and objects in each frame. The resulting representations capture both individual entity states and their cross-entity context. Based on these representations, Target Geometry-Guided Relational Interaction Prediction (TGRIP) predicts target entity interaction states from historical relational memory and target-frame geometry without explicit identity tracking. Motion-Interaction Reconstruction and Alignment (MIRA) then evaluates these predictions through conditional flow reconstruction and semantic consistency checking, providing complementary anomaly evidence beyond prediction error alone. INTERACT achieves state-of-the-art performance, obtaining a frame-level AUC of 84.5% on the ShanghaiTech benchmark. Ablation studies show that removing cross-entity attention causes the largest performance drop, demonstrating the necessity of relational modeling. INTERACT is particularly effective for anomalies caused by changes in relations among people and objects, while maintaining competitive performance in general scenarios. The source code will be made publicly available at https://github.com/ppworkhard/INTERACT.