Papers Text to Audio Retrieval
“Text to Audio Retrieval” 태그가 달린 논문 20편 · 필터 해제
M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP
Contrastive language-audio pre-training (CLAP) has addressed audio-language tasks such as audio-text retrieval by aligning audio and text in a common feature space. While CLAP addresses general audio-language tasks, its …
Audio captioningAudio ClassificationAudio TaggingAudio to Text Retrieval+13Do Audio-Language Models Understand Linguistic Variations?
Open-vocabulary audio language models (ALMs), like Contrastive Language Audio Pretraining (CLAP), represent a promising new paradigm for audio-text retrieval using natural language queries. In this paper, for the first t…
Contrastive LearningNatural Language QueriesRetrievalText Retrieval+1The language of sound search: Examining User Queries in Audio Search Engines
This study examines textual, user-written search queries within the context of sound search engines, encompassing various applications such as foley, sound effects, and general audio retrieval. Current research inadequat…
RetrievalSurveyText to Audio RetrievalEvaluation of pretrained language models on music understanding
Music-text multimodal systems have enabled new approaches to Music Information Research (MIR) applications such as audio-to-text and text-to-audio retrieval, text-based song generation, and music captioning. Despite the …
Music CaptioningNegationSensitivityText to Audio Retrieval+1Dissecting Temporal Understanding in Text-to-Audio Retrieval
Recent advancements in machine learning have fueled research on multimodal tasks, such as for instance text-to-video and text-to-audio retrieval. These tasks require models to understand the semantic content of video and…
AudioCapsRetrievalText to Audio RetrievalEstimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval
Dual-encoder-based audio retrieval systems are commonly optimized with contrastive learning on a set of matching and mismatching audio-caption pairs. This leads to a shared embedding space in which corresponding items fr…
AudioCapsContrastive LearningRetrievalText to Audio RetrievalInternVideo2: Scaling Foundation Models for Multimodal Video Understanding
We introduce InternVideo2, a new family of video foundation models (ViFM) that achieve the state-of-the-art results in video recognition, video-text tasks, and video-centric dialogue. Our core design is a progressive tra…
Action ClassificationAction RecognitionAudio ClassificationContrastive Learning+13WikiMuTe: A web-sourced dataset of semantic descriptions for music audio
Multi-modal deep learning techniques for matching free-form text with music have shown promising results in the field of Music Information Retrieval (MIR). Prior work is often based on large proprietary data while public…
ArticlesCross-Modal RetrievalInformation RetrievalMusic Auto-Tagging+3The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation
We introduce the Song Describer dataset (SDD), a new crowdsourced corpus of high-quality audio-caption pairs, designed for the evaluation of music-and-language models. The dataset consists of 1.1k human-written natural l…
Music CaptioningMusic GenerationRetrievalText to Audio Retrieval+1Advancing Natural-Language Based Audio Retrieval with PaSST and Large Audio-Caption Data Sets
This work presents a text-to-audio-retrieval system based on pre-trained text and spectrogram transformers. Our method projects recordings and textual descriptions into a shared audio-caption space in which related examp…
RetrievalText to Audio RetrievalVAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset
Vision and text have been fully explored in contemporary video-text foundational models, while other modalities such as audio and subtitles in videos have not received sufficient attention. In this paper, we resort to es…
Audio captioningAudio-Visual CaptioningAudio-visual Question AnsweringCross-Modal Retrieval+12ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
In this work, we explore a scalable way for building a general representation model toward unlimited modalities. We release ONE-PEACE, a highly extensible model with 4B parameters that can seamlessly align and integrate …
1 Image, 2*2 StitchiAction ClassificationAudioCapsAudio Classification+18VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset
In this paper, we propose a Vision-Audio-Language Omni-peRception pretraining model (VALOR) for multi-modal understanding and generation. Different from widely-studied vision-language pretraining models, VALOR jointly mo…
Audio captioningAudio-Video Question Answering (AVQA)Audio-Visual CaptioningAudio-visual Question Answering+16Data leakage in cross-modal retrieval training: A case study
The recent progress in text-based audio retrieval was largely propelled by the release of suitable datasets. Since the manual creation of such datasets is a laborious task, obtaining data from online resources can be a c…
Cross-Modal RetrievalRetrievalText to Audio RetrievalExploring Train and Test-Time Augmentations for Audio-Language Learning
In this paper, we aim to unveil the impact of data augmentation in audio-language multi-modal learning, which has not been explored despite its importance. We explore various augmentation methods at not only train-time b…
Audio captioningAudio to Text RetrievalData AugmentationRetrieval+2Matching Text and Audio Embeddings: Exploring Transfer-learning Strategies for Language-based Audio Retrieval
We present an analysis of large-scale pretrained deep learning models used for cross-modal (text-to-audio) retrieval. We use embeddings extracted by these models in a metric learning framework to connect matching pairs o…
Metric LearningRetrievalText to Audio RetrievalTransfer LearningCross Modal Retrieval with Querybank Normalisation
Profiting from large-scale training datasets, advances in neural architecture design and efficient inference, joint embeddings have become the dominant approach for tackling cross-modal retrieval. In this work we first s…
Cross-Modal RetrievalMetric LearningRetrievalText to Audio Retrieval+1Audio Retrieval with Natural Language Queries: A Benchmark Study
The objectives of this work are cross-modal text-audio and audio-text retrieval, in which the goal is to retrieve the audio content from a pool of candidates that best matches a given written description and vice versa. …
AudioCapsAudio captioningAudio to Text RetrievalNatural Language Queries+3OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation
In this paper, we propose an Omni-perception Pre-Trainer (OPT) for cross-modal understanding and generation, by jointly modeling visual, text and audio resources. OPT is constructed in an encoder-decoder framework, inclu…
Audio to Text RetrievalCross-Modal RetrievalDecoderImage Retrieval+2Audio Retrieval with Natural Language Queries
We consider the task of retrieving audio using free-form natural language queries. To study this problem, which has received limited attention in the existing literature, we introduce challenging new benchmarks for text-…
AudioCapsAudio to Text RetrievalAudio/Video to Text RetrievalForm+4