Finetuning Strategies for Querying Sounds by Vocal Imitation

  • 2026-08-19 17:51:50
  • Aditya Bhattacharjee, Christos Plachouras, Sungkyun Chang, Emmanouil Benetos
  • 0

Abstract

This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We investigate two complementary fine-tuning strategies: contrastive learning with a frozen, pretrained CED encoder, and joint contrastive-triplet learning with semi-hard negatives using a MobileNetV3 encoder. This report has been updated for posterity to include details released after the challenge.

 

Quick Read (beta)

loading the full paper ...