arXiv · 2509.10391
Improving Audio Event Recognition with Consistency Regularization
Abstract
Consistency regularization (CR), which enforces agreement between model predictions on augmented views, has found recent benefits in automatic speech recognition [1]. In this paper, we propose the use of consistency regularization for audio event recognition, and demonstrate its effectiveness on AudioSet. With extensive ablation studies for both small ($\sim$20k) and large ($\sim$1.8M) supervised training sets, we show that CR brings consistent improvement over supervised baselines which already heavily utilize data augmentation, and CR using stronger augmentation and multiple augmentations leads to additional gain for the small training set. Furthermore, we extend the use of CR into the semi-supervised setup with 20K labeled samples and 1.8M unlabeled samples, and obtain performance improvement over our best model trained on the small set.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shanmuka Sadhu, Weiran Wang. 2025-09-12. Improving Audio Event Recognition with Consistency Regularization. https://arxiv.org/abs/2509.10391
Cite the original work for its findings. Save a collection to share your selection of sources.