Text to Audio Retrieval
4개 벤치마크 · 논문 20편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
Evaluation of pretrained language models on music understanding
Papers
M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP
Contrastive language-audio pre-training (CLAP) has addressed audio-language tasks such as audio-text retrieval by aligning audio and text in a common feature space. While CLAP addresses general audio-language tasks, its …
Audio captioningAudio ClassificationAudio TaggingAudio to Text Retrieval+13Do Audio-Language Models Understand Linguistic Variations?
Open-vocabulary audio language models (ALMs), like Contrastive Language Audio Pretraining (CLAP), represent a promising new paradigm for audio-text retrieval using natural language queries. In this paper, for the first t…
Contrastive LearningNatural Language QueriesRetrievalText Retrieval+1The language of sound search: Examining User Queries in Audio Search Engines
This study examines textual, user-written search queries within the context of sound search engines, encompassing various applications such as foley, sound effects, and general audio retrieval. Current research inadequat…
RetrievalSurveyText to Audio RetrievalEvaluation of pretrained language models on music understanding
Music-text multimodal systems have enabled new approaches to Music Information Research (MIR) applications such as audio-to-text and text-to-audio retrieval, text-based song generation, and music captioning. Despite the …
Music CaptioningNegationSensitivityText to Audio Retrieval+1Dissecting Temporal Understanding in Text-to-Audio Retrieval
Recent advancements in machine learning have fueled research on multimodal tasks, such as for instance text-to-video and text-to-audio retrieval. These tasks require models to understand the semantic content of video and…
AudioCapsRetrievalText to Audio RetrievalEstimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval
Dual-encoder-based audio retrieval systems are commonly optimized with contrastive learning on a set of matching and mismatching audio-caption pairs. This leads to a shared embedding space in which corresponding items fr…
AudioCapsContrastive LearningRetrievalText to Audio Retrieval