SearcharxivSearch

arXiv subjects

Benfeng Wang

Publications and source records attributed to Benfeng Wang.

3 recordsLinked to original sources

Advancing Adaptive Multi-Stage Video Anomaly Reasoning: A Benchmark Dataset and Method

Recent progress in reasoning capabilities of Multimodal Large Language Models(MLLMs) has highlighted their potential for performing complex video understanding tasks. However, in the domain of Video Anomaly Detection and Understanding (VAD&U), existing MLLM-based methods are largely limited to anomaly localization or post-hoc description, lacking explicit reasoning processes, risk awareness, and decision-oriented interpretation. To address this gap, we define a new task termed Video Anomaly Reasoning (VAR), which elevates video anomaly analysis from descriptive understanding to structured, multi-stage reasoning. VAR explicitly requires models to perform progressive reasoning over anomalous events before answering anomaly-related questions, encompassing visual perception, causal interpretation, and risk-aware decision making. To support this task, we present a new dataset with 8,641 videos, where each video is annotated with diverse question types corresponding to different reasoning depths, totaling more than 50,000 samples, making it one of the largest datasets for video anomaly. The annotations are based on a structured Perception-Cognition-Action Chain-of-Thought (PerCoAct-CoT), which formalizes domain-specific reasoning priors for video anomaly understanding. This design enables systematic evaluation of multi-stage and adaptive anomaly reasoning. In addition, we propose Anomaly-Aware Group Relative Policy Optimization to further enhance reasoning reliability under weak supervision. Building upon the proposed task and dataset, we develop an end-to-end MLLM-based VAR model termed Vad-R1-Plus, which supports adaptive hierarchical reasoning and risk-aware decision making. Extensive experiments demonstrate that the proposed benchmark and method effectively advance the reasoning capabilities of MLLMs on VAR tasks, outperforming both open-source and proprietary baselines.

cs.CV

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought

Recent advancements in reasoning capability of Multimodal Large Language Models (MLLMs) demonstrate its effectiveness in tackling complex visual tasks. However, existing MLLM-based Video Anomaly Detection (VAD) methods remain limited to shallow anomaly descriptions without deep reasoning. In this paper, we propose a new task named Video Anomaly Reasoning (VAR), which aims to enable deep analysis and understanding of anomalies in the video by requiring MLLMs to think explicitly before answering. To this end, we propose Vad-R1, an end-to-end MLLM-based framework for VAR. Specifically, we design a Perception-to-Cognition Chain-of-Thought (P2C-CoT) that simulates the human process of recognizing anomalies, guiding the MLLM to reason anomaly step-by-step. Based on the structured P2C-CoT, we construct Vad-Reasoning, a dedicated dataset for VAR. Furthermore, we propose an improved reinforcement learning algorithm AVA-GRPO, which explicitly incentivizes the anomaly reasoning capability of MLLMs through a self-verification mechanism with limited annotations. Experimental results demonstrate that Vad-R1 achieves superior performance, outperforming both open-source and proprietary models on VAD and VAR tasks. Codes and datasets will be released at https://github.com/wbfwonderful/Vad-R1.

cs.CV

A new radial basis function collocation method based on the quasi-uniform nodes for 2D fractional wave equation

We mainly concerned with a decoupled fractional Laplacian wave equation in this paper. A new time-space domain radial basis function (RBF) collocation method is introduced to solve the fractional wave equation, which describes seismic wave propagation in attenuation media. The directional fractional Laplacian is adopted to cope with the fractional Laplacian of RBFs. In order to increase the computational efficiency, we introduced the quasi-uniform nodes configuration scheme, which is suitable for mesh-free discretization of wave equations. The comparison between the new method and the commonly-used pseudo-spectral method are implemented on square homogeneous models with different model size. The CPU time and relative errors of different methods show that the quasi-uniform configuration scheme provides better performance and the calculation efficiency advantage is significantly prominent as the computation domain increases. The relative errors indicate that the RBF collocation method with quasi-uniform configuration could improve the computational efficiency effectively and provide satisfactory accuracy. This advantage was especially highlighted in complex models, where the new approach achieved the same accuracy with only a half number of points. The implementation on the 2D complex model further demonstrated the accuracy, efficiency, and flexibility of the proposed new method.

physics.comp-ph