paper-with-me

Papers Audio to Text Retrieval

“Audio to Text Retrieval” 태그가 달린 논문 10편 · 필터 해제

M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP

2025-03-28 · Daisuke Niizumi, Daiki Takeuchi, Masahiro Yasuda, Binh Thien Nguyen 외

Contrastive language-audio pre-training (CLAP) has addressed audio-language tasks such as audio-text retrieval by aligning audio and text in a common feature space. While CLAP addresses general audio-language tasks, its …

Audio captioningAudio ClassificationAudio TaggingAudio to Text Retrieval+13

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models

2024-11-28 · Jisheng Bai, Haohe Liu, Mou Wang, Dongyuan Shi 외

With the emergence of audio-language models, constructing large-scale paired audio-language datasets has become essential yet challenging for model development, primarily due to the time-intensive and labour-heavy demand…

Audio captioningAudio to Text RetrievalCaption GenerationRetrieval+1

Killing two birds with one stone: Can an audio captioning system also be used for audio-text retrieval?

2023-08-29 · Etienne Labbé, Thomas Pellegrini, Julien Pinquier

Automated Audio Captioning (AAC) aims to develop systems capable of describing an audio recording using a textual sentence. In contrast, Audio-Text Retrieval (ATR) systems seek to find the best matching audio recording(s…

AudioCapsAudio captioningAudio TaggingAudio to Text Retrieval+4

ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities

2023-05-18 · Peng Wang, Shijie Wang, Junyang Lin, Shuai Bai 외

In this work, we explore a scalable way for building a general representation model toward unlimited modalities. We release ONE-PEACE, a highly extensible model with 4B parameters that can seamlessly align and integrate …

1 Image, 2*2 StitchiAction ClassificationAudioCapsAudio Classification+18

On Negative Sampling for Contrastive Audio-Text Retrieval

2022-11-08 · Huang Xie, Okko Räsänen, Tuomas Virtanen

This paper investigates negative sampling for contrastive learning in the context of audio-text retrieval. The strategy for negative sampling refers to selecting negatives (either audio clips or textual descriptions) fro…

Audio to Text RetrievalContrastive LearningRetrievalText Retrieval

Exploring Train and Test-Time Augmentations for Audio-Language Learning

2022-10-31 · Eungbeom Kim, Jinhee Kim, Yoori Oh, KyungSu Kim 외

In this paper, we aim to unveil the impact of data augmentation in audio-language multi-modal learning, which has not been explored despite its importance. We explore various augmentation methods at not only train-time b…

Audio captioningAudio to Text RetrievalData AugmentationRetrieval+2

Contrastive Audio-Language Learning for Music

2022-08-25 · Ilaria Manco, Emmanouil Benetos, Elio Quinton, György Fazekas

As one of the most intuitive interfaces known to humans, natural language has the potential to mediate many tasks that involve human-computer interaction, especially in application-focused fields like Music Information R…

Audio to Text RetrievalDescriptiveGenre classificationInformation Retrieval+3

Audio Retrieval with Natural Language Queries: A Benchmark Study

2021-12-17 · A. Sophia Koepke, Andreea-Maria Oncescu, João F. Henriques, Zeynep Akata 외

The objectives of this work are cross-modal text-audio and audio-text retrieval, in which the goal is to retrieve the audio content from a pool of candidates that best matches a given written description and vice versa. …

AudioCapsAudio captioningAudio to Text RetrievalNatural Language Queries+3

OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation

2021-07-01 · Jing Liu, Xinxin Zhu, Fei Liu, Longteng Guo 외

In this paper, we propose an Omni-perception Pre-Trainer (OPT) for cross-modal understanding and generation, by jointly modeling visual, text and audio resources. OPT is constructed in an encoder-decoder framework, inclu…

Audio to Text RetrievalCross-Modal RetrievalDecoderImage Retrieval+2

Audio Retrieval with Natural Language Queries

2021-05-05 · Andreea-Maria Oncescu, A. Sophia Koepke, João F. Henriques, Zeynep Akata 외

We consider the task of retrieving audio using free-form natural language queries. To study this problem, which has received limited attention in the existing literature, we introduce challenging new benchmarks for text-…

AudioCapsAudio to Text RetrievalAudio/Video to Text RetrievalForm+4
1–10 / 10