arXiv · 2402.01274
On the Transferability of Large-Scale Self-Supervision to Few-Shot Audio Classification
Abstract
In recent years, self-supervised learning has excelled for its capacity to learn robust feature representations from unlabelled data. Networks pretrained through self-supervision serve as effective feature extractors for downstream tasks, including Few-Shot Learning. While the evaluation of unsupervised approaches for few-shot learning is well-established in imagery, it is notably absent in acoustics. This study addresses this gap by assessing large-scale self-supervised models' performance in few-shot audio classification. Additionally, we explore the relationship between a model's few-shot learning capability and other downstream task benchmarks. Our findings reveal state-of-the-art performance in some few-shot problems such as SpeechCommandsv2, as well as strong correlations between speech-based few-shot problems and various downstream audio tasks.
Explore related subjects
Keep this discovery
Calum Heggan, Sam Budgett, Timothy Hospedales, Mehrdad Yaghoobi. 2024-02-02. On the Transferability of Large-Scale Self-Supervision to Few-Shot Audio Classification. https://arxiv.org/abs/2402.01274
Cite the original work for its findings. Save a collection to share your selection of sources.