Papers Audio to Text Retrieval
“Audio to Text Retrieval” 태그가 달린 논문 10편 · 필터 해제
M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP
Contrastive language-audio pre-training (CLAP) has addressed audio-language tasks such as audio-text retrieval by aligning audio and text in a common feature space. While CLAP addresses general audio-language tasks, its …
Audio captioningAudio ClassificationAudio TaggingAudio to Text Retrieval+13AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models
With the emergence of audio-language models, constructing large-scale paired audio-language datasets has become essential yet challenging for model development, primarily due to the time-intensive and labour-heavy demand…
Audio captioningAudio to Text RetrievalCaption GenerationRetrieval+1Killing two birds with one stone: Can an audio captioning system also be used for audio-text retrieval?
Automated Audio Captioning (AAC) aims to develop systems capable of describing an audio recording using a textual sentence. In contrast, Audio-Text Retrieval (ATR) systems seek to find the best matching audio recording(s…
AudioCapsAudio captioningAudio TaggingAudio to Text Retrieval+4ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
In this work, we explore a scalable way for building a general representation model toward unlimited modalities. We release ONE-PEACE, a highly extensible model with 4B parameters that can seamlessly align and integrate …
1 Image, 2*2 StitchiAction ClassificationAudioCapsAudio Classification+18On Negative Sampling for Contrastive Audio-Text Retrieval
This paper investigates negative sampling for contrastive learning in the context of audio-text retrieval. The strategy for negative sampling refers to selecting negatives (either audio clips or textual descriptions) fro…
Audio to Text RetrievalContrastive LearningRetrievalText RetrievalExploring Train and Test-Time Augmentations for Audio-Language Learning
In this paper, we aim to unveil the impact of data augmentation in audio-language multi-modal learning, which has not been explored despite its importance. We explore various augmentation methods at not only train-time b…
Audio captioningAudio to Text RetrievalData AugmentationRetrieval+2Contrastive Audio-Language Learning for Music
As one of the most intuitive interfaces known to humans, natural language has the potential to mediate many tasks that involve human-computer interaction, especially in application-focused fields like Music Information R…
Audio to Text RetrievalDescriptiveGenre classificationInformation Retrieval+3Audio Retrieval with Natural Language Queries: A Benchmark Study
The objectives of this work are cross-modal text-audio and audio-text retrieval, in which the goal is to retrieve the audio content from a pool of candidates that best matches a given written description and vice versa. …
AudioCapsAudio captioningAudio to Text RetrievalNatural Language Queries+3OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation
In this paper, we propose an Omni-perception Pre-Trainer (OPT) for cross-modal understanding and generation, by jointly modeling visual, text and audio resources. OPT is constructed in an encoder-decoder framework, inclu…
Audio to Text RetrievalCross-Modal RetrievalDecoderImage Retrieval+2Audio Retrieval with Natural Language Queries
We consider the task of retrieving audio using free-form natural language queries. To study this problem, which has received limited attention in the existing literature, we introduce challenging new benchmarks for text-…
AudioCapsAudio to Text RetrievalAudio/Video to Text RetrievalForm+4