paper-with-me

Text to Audio Retrieval

4개 벤치마크 · 논문 20편 · 이 태스크의 논문 보기 →

Benchmarks

Clotho

결과 12개

AudioCaps

결과 11개

SoundDescs

결과 4개

Localized Narratives

결과 1개

Most implemented

Papers

M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP

2025-03-28 · Daisuke Niizumi, Daiki Takeuchi, Masahiro Yasuda, Binh Thien Nguyen 외

Contrastive language-audio pre-training (CLAP) has addressed audio-language tasks such as audio-text retrieval by aligning audio and text in a common feature space. While CLAP addresses general audio-language tasks, its …

Audio captioningAudio ClassificationAudio TaggingAudio to Text Retrieval+13

Do Audio-Language Models Understand Linguistic Variations?

2024-10-21 · Ramaneswaran Selvakumar, Sonal Kumar, Hemant Kumar Giri, Nishit Anand 외

Open-vocabulary audio language models (ALMs), like Contrastive Language Audio Pretraining (CLAP), represent a promising new paradigm for audio-text retrieval using natural language queries. In this paper, for the first t…

Contrastive LearningNatural Language QueriesRetrievalText Retrieval+1

The language of sound search: Examining User Queries in Audio Search Engines

2024-10-10 · Benno Weck, Frederic Font

This study examines textual, user-written search queries within the context of sound search engines, encompassing various applications such as foley, sound effects, and general audio retrieval. Current research inadequat…

RetrievalSurveyText to Audio Retrieval

Evaluation of pretrained language models on music understanding

2024-09-17 · Yannis Vasilakis, Rachel Bittner, Johan Pauwels

Music-text multimodal systems have enabled new approaches to Music Information Research (MIR) applications such as audio-to-text and text-to-audio retrieval, text-based song generation, and music captioning. Despite the …

Music CaptioningNegationSensitivityText to Audio Retrieval+1

Dissecting Temporal Understanding in Text-to-Audio Retrieval

2024-09-01 · Andreea-Maria Oncescu, João F. Henriques, A. Sophia Koepke

Recent advancements in machine learning have fueled research on multimodal tasks, such as for instance text-to-video and text-to-audio retrieval. These tasks require models to understand the semantic content of video and…

AudioCapsRetrievalText to Audio Retrieval

Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval

2024-08-21 · Paul Primus, Florian Schmid, Gerhard Widmer

Dual-encoder-based audio retrieval systems are commonly optimized with contrastive learning on a set of matching and mismatching audio-caption pairs. This leads to a shared embedding space in which corresponding items fr…

AudioCapsContrastive LearningRetrievalText to Audio Retrieval

전체 20편 보기 →