paper-with-me

홈 › Papers

Retrieval-Augmented Text-to-Audio Generation

2023-09-14 · Yi Yuan, Haohe Liu, Xubo Liu, Qiushi Huang, Mark D. Plumbley, Wenwu Wang

Despite recent progress in text-to-audio (TTA) generation, we show that the state-of-the-art models, such as AudioLDM, trained on datasets with an imbalanced class distribution, such as AudioCaps, are biased in their generation performance. Specifically, they excel in generating common audio classes while underperforming in the rare ones, thus degrading the overall generation performance. We refer to this problem as long-tailed text-to-audio generation. To address this issue, we propose a simple retrieval-augmented approach for TTA models. Specifically, given an input text prompt, we first leverage a Contrastive Language Audio Pretraining (CLAP) model to retrieve relevant text-audio pairs. The features of the retrieved audio-text data are then used as additional conditions to guide the learning of TTA models. We enhance AudioLDM with our proposed approach and denote the resulting augmented system as Re-AudioLDM. On the AudioCaps dataset, Re-AudioLDM achieves a state-of-the-art Frechet Audio Distance (FAD) of 1.37, outperforming the existing approaches by a large margin. Furthermore, we show that Re-AudioLDM can generate realistic audio for complex scenes, rare audio classes, and even unseen audio types, indicating its potential in TTA tasks.

📄 PDF Abstract BibTeX arXiv:2309.08051

Code (0)

등록된 구현이 없습니다.

Tasks

AudioCapsAudio GenerationFADRetrieval

Similar Papers 제목 키워드 기반

Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation

2024-11-07 · Mu Yang, Bowen Shi, Matthew Le, Wei-Ning Hsu 외

This work focuses on improving Text-To-Audio (TTA) generation on zero-shot and few-shot settings (i.e. generating unseen or uncommon audio events). Inspired by the success of Retrieval-Augmented Generation (RAG) in Large…

Audio GenerationLarge Language ModelRAGRetrieval+1

WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models

2025-02-20 · Yifu Chen, Shengpeng Ji, Haoxiao Wang, Ziqing Wang 외

Retrieval Augmented Generation (RAG) has gained widespread adoption owing to its capacity to empower large language models (LLMs) to integrate external knowledge. However, existing RAG frameworks are primarily designed f…

Automatic Speech RecognitionRAGRetrievalRetrieval-augmented Generation+2

Enhancing Retrieval-Augmented Audio Captioning with Generation-Assisted Multimodal Querying and Progressive Learning

2024-10-14 · Choi Changin, Lim Sungjun, Rhee Wonjong

Retrieval-augmented generation can improve audio captioning by incorporating relevant audio-text pairs from a knowledge base. Existing methods typically rely solely on the input audio as a unimodal retrieval query. In co…

AudioCapsAudio captioningRAGRetrieval+1

Speech Retrieval-Augmented Generation without Automatic Speech Recognition

2024-12-21 · Do June Min, Karel Mundnich, Andy Lapastora, Erfan Soltanmohammadi 외

One common approach for question answering over speech data is to first transcribe speech using automatic speech recognition (ASR) and then employ text-based retrieval-augmented generation (RAG) on the transcriptions. Wh…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+7

Retrieval-Augmented Audio Deepfake Detection

2024-04-22 · Zuheng Kang, Yayun He, Botao Zhao, Xiaoyang Qu 외

With recent advances in speech synthesis including text-to-speech (TTS) and voice conversion (VC) systems enabling the generation of ultra-realistic audio deepfakes, there is growing concern about their potential misuse.…

Audio Deepfake DetectionDeepFake DetectionFace SwappingRAG+6