paper-with-me

홈 › Papers

Improving Natural-Language-based Audio Retrieval with Transfer Learning and Audio & Text Augmentations

2022-08-24 · Paul Primus, Gerhard Widmer

The absence of large labeled datasets remains a significant challenge in many application areas of deep learning. Researchers and practitioners typically resort to transfer learning and data augmentation to alleviate this issue. We study these strategies in the context of audio retrieval with natural language queries (Task 6b of the DCASE 2022 Challenge). Our proposed system uses pre-trained embedding models to project recordings and textual descriptions into a shared audio-caption space in which related examples from different modalities are close. We employ various data augmentation techniques on audio and text inputs and systematically tune their corresponding hyperparameters with sequential model-based optimization. Our results show that the used augmentations strategies reduce overfitting and improve retrieval performance.

📄 PDF Abstract BibTeX arXiv:2208.11460

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationNatural Language QueriesRetrievalTransfer Learning

Similar Papers 제목 키워드 기반

ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models

2026-06-27 · Fengjie Lu, Chenang Jiang, Jiarui Hai, Helin Wang 외 arxiv

Recent advances in language--audio retrieval have been largely driven by contrastive dual-encoder architectures that align audio and text in a shared embedding space. While effective, existing retrieval embeddings are pr…

Question Answering

Contrastive Audio-Language Learning for Music

2022-08-25 · Ilaria Manco, Emmanouil Benetos, Elio Quinton, György Fazekas

As one of the most intuitive interfaces known to humans, natural language has the potential to mediate many tasks that involve human-computer interaction, especially in application-focused fields like Music Information R…

Audio to Text RetrievalDescriptiveGenre classificationInformation Retrieval+3

Audio Retrieval with Natural Language Queries: A Benchmark Study

2021-12-17 · A. Sophia Koepke, Andreea-Maria Oncescu, João F. Henriques, Zeynep Akata 외

The objectives of this work are cross-modal text-audio and audio-text retrieval, in which the goal is to retrieve the audio content from a pool of candidates that best matches a given written description and vice versa. …

AudioCapsAudio captioningAudio to Text RetrievalNatural Language Queries+3

Audio Retrieval with WavText5K and CLAP Training

2022-09-28 · Soham Deshmukh, Benjamin Elizalde, Huaming Wang

Audio-Text retrieval takes a natural language query to retrieve relevant audio files in a database. Conversely, Text-Audio retrieval takes an audio file as a query to retrieve relevant natural language descriptions. Most…

AudioCapsAudio captioningContrastive LearningRetrieval+1

Audio Retrieval with Natural Language Queries

2021-05-05 · Andreea-Maria Oncescu, A. Sophia Koepke, João F. Henriques, Zeynep Akata 외

We consider the task of retrieving audio using free-form natural language queries. To study this problem, which has received limited attention in the existing literature, we introduce challenging new benchmarks for text-…

AudioCapsAudio to Text RetrievalAudio/Video to Text RetrievalForm+4