arXiv · 2110.08895
DECAR: Deep Clustering for learning general-purpose Audio Representations
Abstract
We introduce DECAR, a self-supervised pre-training approach for learning general-purpose audio representations. Our system is based on clustering: it utilizes an offline clustering step to provide target labels that act as pseudo-labels for solving a prediction task. We develop on top of recent advances in self-supervised learning for computer vision and design a lightweight, easy-to-use self-supervised pre-training scheme. We pre-train DECAR embeddings on a balanced subset of the large-scale Audioset dataset and transfer those representations to 9 downstream classification tasks, including speech, music, animal sounds, and acoustic scenes. Furthermore, we conduct ablation studies identifying key design choices and also make all our code and pre-trained models publicly available.
Explore related subjects
Keep this discovery
Sreyan Ghosh, Sandesh V Katta, Ashish Seth, S. Umesh. 2021-10-17. DECAR: Deep Clustering for learning general-purpose Audio Representations. https://doi.org/10.1109/jstsp.2022.3202093
Cite the original work for its findings. Save a collection to share your selection of sources.