paper-with-me

Papers

Finetuning Strategies for Querying Sounds by Vocal Imitation

2026-08-19 · Aditya Bhattacharjee, Christos Plachouras, Sungkyun Chang, Emmanouil Benetos arxiv

This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We investigate two complementary fine-tuning strategies: contrastive learning with a frozen, pretrained CED encoder, and joint contrastive-triplet learning with semi-hard negatives using a MobileNetV3 encoder. This report has been updated for posterity to include details released after the challenge.

📄 PDF Abstract BibTeX arXiv:2608.19174

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive Learning

Similar Papers 제목 키워드 기반

Sketching With Your Voice: "Non-Phonorealistic" Rendering of Sounds via Vocal Imitation

2024-09-20 · Matthew Caren, Kartik Chandra, Joshua B. Tenenbaum, Jonathan Ragan-Kelley 외

We present a method for automatically producing human-like vocal imitations of sounds: the equivalent of "sketching," but for auditory rather than visual representation. Starting with a simulated model of the human vocal…

Deep learning for detection of bird vocalisations

2016-09-25 · Ilyas Potamitis

This work focuses on reliable detection of bird sound emissions as recorded in the open field. Acoustic detection of avian sounds can be used for the automatized monitoring of multiple bird taxa and querying in long-term…

Deep Learning

Low-dimensional representation of infant and adult vocalization acoustics

2022-04-25 · Silvia Pagliarini, Sara Schneider, Christopher T. Kello, Anne S. Warlaumont

During the first years of life, infant vocalizations change considerably, as infants develop the vocalization skills that enable them to produce speech sounds. Characterizations based on specific acoustic features, proto…

Vocalsound: A Dataset for Improving Human Vocal Sounds Recognition

2022-05-06 · Yuan Gong, Jin Yu, James Glass

Recognizing human non-speech vocalizations is an important task and has broad applications such as automatic sound transcription and health condition monitoring. However, existing datasets have a relatively small number …

Audio Classification

Vocal Breath Sound Based Gender Classification

2022-11-11 · Mohammad Shaique Solanki, Ashutosh M Bharadwaj, Jeevan K, Prasanta Kumar Ghosh

Voiced speech signals such as continuous speech are known to have acoustic features such as pitch(F0), and formant frequencies(F1, F2, F3) which can be used for gender classification. However, gender classification studi…

ClassificationGender Classification