arXiv · 2608.19174
Finetuning Strategies for Querying Sounds by Vocal Imitation
Abstract
This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We investigate two complementary fine-tuning strategies: contrastive learning with a frozen, pretrained CED encoder, and joint contrastive-triplet learning with semi-hard negatives using a MobileNetV3 encoder. This report has been updated for posterity to include details released after the challenge.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Aditya Bhattacharjee, Christos Plachouras, Sungkyun Chang, Emmanouil Benetos. 2026-08-19. Finetuning Strategies for Querying Sounds by Vocal Imitation. https://arxiv.org/abs/2608.19174
Cite the original work for its findings. Save a collection to share your selection of sources.