arXiv · 2603.04605
Temporal Pooling Strategies for Training-Free Anomalous Sound Detection with Self-Supervised Audio Embeddings
Abstract
Training-free anomalous sound detection (ASD) based on pre-trained audio embedding models has recently garnered significant attention, as it enables the detection of anomalous sounds using only normal reference data without task-specific model training or fine-tuning. However, existing embedding-based approaches almost exclusively rely on temporal mean pooling, leaving temporal pooling in training-free ASD largely unexplored. In this paper, we present the first systematic evaluation of temporal pooling strategies for training-free ASD with pre-trained audio embeddings. We propose relative deviation pooling (RDP), an adaptive pooling method that assigns larger weights to embeddings with stronger temporal deviations, investigate feature-wise non-linear aggregation using generalized mean (GeM) pooling, and examine a hybrid combination of both strategies. Experiments on five benchmark datasets demonstrate that the proposed pooling strategies consistently outperform mean pooling and achieve state-of-the-art performance for training-free ASD, including results that surpass previously reported trained systems and ensembles on the DCASE2025 ASD dataset.
Explore related subjects
Keep this discovery
Kevin Wilkinghoff, Sarthak Yadav, Zheng-Hua Tan. 2026-03-04. Temporal Pooling Strategies for Training-Free Anomalous Sound Detection with Self-Supervised Audio Embeddings. https://arxiv.org/abs/2603.04605
Cite the original work for its findings. Save a collection to share your selection of sources.