arXiv · 2609.22136
DiFA: Dual Evidence Fusion and Aggregation for Token-Level Text Anomaly Detection
Abstract
Text anomaly detection, the task of identifying text instances that deviate from normal language patterns, is crucial for language-driven applications. However, most existing methods can only perform document-level anomaly detection, making it hard to locate harmful phrases or support targeted prevention. Recently, there has been an emerging trend toward token-level text anomaly detection, which aims to address the above limitation by identifying anomalous words or fragments within a document. Nevertheless, one representative method mainly relies on representation-space distance measurement, neglecting the complementary roles of different anomaly cues in capturing diverse abnormal patterns. To bridge the gaps, we propose a Dual-evidence framework with adaptive Fusion and Aggregation (DiFA) for token-level anomaly detection. DiFA derives anomaly scores from form-structural and semantic views to capture visible structural abnormality and contextual inconsistency, respectively, thereby providing complementary evidence for identifying diverse anomalies. To combine these two scores with varying numerical scales, DiFA incorporates a calibration and fusion mechanism to adaptively balance the two views. Moreover, to obtain a discriminative document-level score, a multivariate aggregation method is designed to summarize token-level anomaly scores from multiple perspectives, preventing rare anomalous tokens from being diluted. Extensive experiments across various text anomaly detection benchmarks demonstrate that DiFA consistently achieves top performance while maintaining strong efficiency, robustness, and interpretability. The code and scripts are available at: https://github.com/qyy11-com/DiFA.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yanyu Qian, Pengcheng Weng, Yue Tan, Enguang Zuo, Yu Zheng, Yixin Liu. 2026-08-24. DiFA: Dual Evidence Fusion and Aggregation for Token-Level Text Anomaly Detection. https://arxiv.org/abs/2609.22136
Cite the original work for its findings. Save a collection to share your selection of sources.