Improving Natural-Language-based Audio Retrieval with Transfer Learning and Audio & Text Augmentations
The absence of large labeled datasets remains a significant challenge in many application areas of deep learning. Researchers and practitioners typically resort to transfer learning and data augmentation to alleviate this issue. We study these strategies in the context of audio retrieval with natural language queries (Task 6b of the DCASE 2022 Challenge). Our proposed system uses pre-trained embedding models to project recordings and textual descriptions into a shared audio-caption space in which related examples from different modalities are close. We employ various data augmentation techniques on audio and text inputs and systematically tune their corresponding hyperparameters with sequential model-based optimization. Our results show that the used augmentations strategies reduce overfitting and improve retrieval performance.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationNatural Language QueriesRetrievalTransfer LearningSimilar Papers 제목 키워드 기반
ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models
Recent advances in language--audio retrieval have been largely driven by contrastive dual-encoder architectures that align audio and text in a shared embedding space. While effective, existing retrieval embeddings are pr…
Question AnsweringContrastive Audio-Language Learning for Music
As one of the most intuitive interfaces known to humans, natural language has the potential to mediate many tasks that involve human-computer interaction, especially in application-focused fields like Music Information R…
Audio to Text RetrievalDescriptiveGenre classificationInformation Retrieval+3Audio Retrieval with Natural Language Queries: A Benchmark Study
The objectives of this work are cross-modal text-audio and audio-text retrieval, in which the goal is to retrieve the audio content from a pool of candidates that best matches a given written description and vice versa. …
AudioCapsAudio captioningAudio to Text RetrievalNatural Language Queries+3Audio Retrieval with WavText5K and CLAP Training
Audio-Text retrieval takes a natural language query to retrieve relevant audio files in a database. Conversely, Text-Audio retrieval takes an audio file as a query to retrieve relevant natural language descriptions. Most…
AudioCapsAudio captioningContrastive LearningRetrieval+1Audio Retrieval with Natural Language Queries
We consider the task of retrieving audio using free-form natural language queries. To study this problem, which has received limited attention in the existing literature, we introduce challenging new benchmarks for text-…
AudioCapsAudio to Text RetrievalAudio/Video to Text RetrievalForm+4