paper-with-me

Papers Text to Audio Retrieval

“Text to Audio Retrieval” 태그가 달린 논문 20편 · 필터 해제

M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP

2025-03-28 · Daisuke Niizumi, Daiki Takeuchi, Masahiro Yasuda, Binh Thien Nguyen 외

Contrastive language-audio pre-training (CLAP) has addressed audio-language tasks such as audio-text retrieval by aligning audio and text in a common feature space. While CLAP addresses general audio-language tasks, its …

Audio captioningAudio ClassificationAudio TaggingAudio to Text Retrieval+13

Do Audio-Language Models Understand Linguistic Variations?

2024-10-21 · Ramaneswaran Selvakumar, Sonal Kumar, Hemant Kumar Giri, Nishit Anand 외

Open-vocabulary audio language models (ALMs), like Contrastive Language Audio Pretraining (CLAP), represent a promising new paradigm for audio-text retrieval using natural language queries. In this paper, for the first t…

Contrastive LearningNatural Language QueriesRetrievalText Retrieval+1

The language of sound search: Examining User Queries in Audio Search Engines

2024-10-10 · Benno Weck, Frederic Font

This study examines textual, user-written search queries within the context of sound search engines, encompassing various applications such as foley, sound effects, and general audio retrieval. Current research inadequat…

RetrievalSurveyText to Audio Retrieval

Evaluation of pretrained language models on music understanding

2024-09-17 · Yannis Vasilakis, Rachel Bittner, Johan Pauwels

Music-text multimodal systems have enabled new approaches to Music Information Research (MIR) applications such as audio-to-text and text-to-audio retrieval, text-based song generation, and music captioning. Despite the …

Music CaptioningNegationSensitivityText to Audio Retrieval+1

Dissecting Temporal Understanding in Text-to-Audio Retrieval

2024-09-01 · Andreea-Maria Oncescu, João F. Henriques, A. Sophia Koepke

Recent advancements in machine learning have fueled research on multimodal tasks, such as for instance text-to-video and text-to-audio retrieval. These tasks require models to understand the semantic content of video and…

AudioCapsRetrievalText to Audio Retrieval

Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval

2024-08-21 · Paul Primus, Florian Schmid, Gerhard Widmer

Dual-encoder-based audio retrieval systems are commonly optimized with contrastive learning on a set of matching and mismatching audio-caption pairs. This leads to a shared embedding space in which corresponding items fr…

AudioCapsContrastive LearningRetrievalText to Audio Retrieval

InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

2024-03-22 · Yi Wang, Kunchang Li, Xinhao Li, Jiashuo Yu 외

We introduce InternVideo2, a new family of video foundation models (ViFM) that achieve the state-of-the-art results in video recognition, video-text tasks, and video-centric dialogue. Our core design is a progressive tra…

Action ClassificationAction RecognitionAudio ClassificationContrastive Learning+13

WikiMuTe: A web-sourced dataset of semantic descriptions for music audio

2023-12-14 · Benno Weck, Holger Kirchhoff, Peter Grosche, Xavier Serra

Multi-modal deep learning techniques for matching free-form text with music have shown promising results in the field of Music Information Retrieval (MIR). Prior work is often based on large proprietary data while public…

ArticlesCross-Modal RetrievalInformation RetrievalMusic Auto-Tagging+3

The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation

2023-11-16 · Ilaria Manco, Benno Weck, Seungheon Doh, Minz Won 외

We introduce the Song Describer dataset (SDD), a new crowdsourced corpus of high-quality audio-caption pairs, designed for the evaluation of music-and-language models. The dataset consists of 1.1k human-written natural l…

Music CaptioningMusic GenerationRetrievalText to Audio Retrieval+1

Advancing Natural-Language Based Audio Retrieval with PaSST and Large Audio-Caption Data Sets

2023-08-08 · Paul Primus, Khaled Koutini, Gerhard Widmer

This work presents a text-to-audio-retrieval system based on pre-trained text and spectrogram transformers. Our method projects recordings and textual descriptions into a shared audio-caption space in which related examp…

RetrievalText to Audio Retrieval

VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset

2023-05-29 · NeurIPS 2023 11 · Sihan Chen, Handong Li, Qunbo Wang, Zijia Zhao 외

Vision and text have been fully explored in contemporary video-text foundational models, while other modalities such as audio and subtitles in videos have not received sufficient attention. In this paper, we resort to es…

Audio captioningAudio-Visual CaptioningAudio-visual Question AnsweringCross-Modal Retrieval+12

ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities

2023-05-18 · Peng Wang, Shijie Wang, Junyang Lin, Shuai Bai 외

In this work, we explore a scalable way for building a general representation model toward unlimited modalities. We release ONE-PEACE, a highly extensible model with 4B parameters that can seamlessly align and integrate …

1 Image, 2*2 StitchiAction ClassificationAudioCapsAudio Classification+18

VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset

2023-04-17 · Jing Liu, Sihan Chen, Xingjian He, Longteng Guo 외

In this paper, we propose a Vision-Audio-Language Omni-peRception pretraining model (VALOR) for multi-modal understanding and generation. Different from widely-studied vision-language pretraining models, VALOR jointly mo…

Audio captioningAudio-Video Question Answering (AVQA)Audio-Visual CaptioningAudio-visual Question Answering+16

Data leakage in cross-modal retrieval training: A case study

2023-02-23 · Benno Weck, Xavier Serra

The recent progress in text-based audio retrieval was largely propelled by the release of suitable datasets. Since the manual creation of such datasets is a laborious task, obtaining data from online resources can be a c…

Cross-Modal RetrievalRetrievalText to Audio Retrieval

Exploring Train and Test-Time Augmentations for Audio-Language Learning

2022-10-31 · Eungbeom Kim, Jinhee Kim, Yoori Oh, KyungSu Kim 외

In this paper, we aim to unveil the impact of data augmentation in audio-language multi-modal learning, which has not been explored despite its importance. We explore various augmentation methods at not only train-time b…

Audio captioningAudio to Text RetrievalData AugmentationRetrieval+2

Matching Text and Audio Embeddings: Exploring Transfer-learning Strategies for Language-based Audio Retrieval

2022-10-06 · Benno Weck, Miguel Pérez Fernández, Holger Kirchhoff, Xavier Serra

We present an analysis of large-scale pretrained deep learning models used for cross-modal (text-to-audio) retrieval. We use embeddings extracted by these models in a metric learning framework to connect matching pairs o…

Metric LearningRetrievalText to Audio RetrievalTransfer Learning

Cross Modal Retrieval with Querybank Normalisation

2021-12-23 · CVPR 2022 1 · Simion-Vlad Bogolin, Ioana Croitoru, Hailin Jin, Yang Liu 외

Profiting from large-scale training datasets, advances in neural architecture design and efficient inference, joint embeddings have become the dominant approach for tackling cross-modal retrieval. In this work we first s…

Cross-Modal RetrievalMetric LearningRetrievalText to Audio Retrieval+1

Audio Retrieval with Natural Language Queries: A Benchmark Study

2021-12-17 · A. Sophia Koepke, Andreea-Maria Oncescu, João F. Henriques, Zeynep Akata 외

The objectives of this work are cross-modal text-audio and audio-text retrieval, in which the goal is to retrieve the audio content from a pool of candidates that best matches a given written description and vice versa. …

AudioCapsAudio captioningAudio to Text RetrievalNatural Language Queries+3

OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation

2021-07-01 · Jing Liu, Xinxin Zhu, Fei Liu, Longteng Guo 외

In this paper, we propose an Omni-perception Pre-Trainer (OPT) for cross-modal understanding and generation, by jointly modeling visual, text and audio resources. OPT is constructed in an encoder-decoder framework, inclu…

Audio to Text RetrievalCross-Modal RetrievalDecoderImage Retrieval+2

Audio Retrieval with Natural Language Queries

2021-05-05 · Andreea-Maria Oncescu, A. Sophia Koepke, João F. Henriques, Zeynep Akata 외

We consider the task of retrieving audio using free-form natural language queries. To study this problem, which has received limited attention in the existing literature, we introduce challenging new benchmarks for text-…

AudioCapsAudio to Text RetrievalAudio/Video to Text RetrievalForm+4
1–20 / 20