paper-with-me

홈 › Papers

Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning

2023-09-20 · Luoyi Sun, Xuenan Xu, Mengyue Wu, Weidi Xie

Recently, the AI community has made significant strides in developing powerful foundation models, driven by large-scale multimodal datasets. However, for audio representation learning, existing datasets suffer from limitations in the following aspects: insufficient volume, simplistic content, and arduous collection procedures. To establish an audio dataset with high-quality captions, we propose an innovative, automatic approach leveraging multimodal inputs, such as video frames, audio streams. Specifically, we construct a large-scale, high-quality, audio-language dataset, named as Auto-ACD, comprising over 1.5M audio-text pairs. We exploit a series of pre-trained models or APIs, to determine audio-visual synchronisation, generate image captions, object detection, or audio tags for specific videos. Subsequently, we employ LLM to paraphrase a congruent caption for each audio, guided by the extracted multi-modality clues. To demonstrate the effectiveness of the proposed dataset, we train widely used models on our dataset and show performance improvement on various downstream tasks, for example, audio-language retrieval, audio captioning, zero-shot classification. In addition, we establish a novel benchmark with environmental information and provide a benchmark for audio-text tasks.

📄 PDF Abstract BibTeX arXiv:2309.11500

Code (0)

등록된 구현이 없습니다.

Tasks

Audio captioningCaption GenerationImage Captioningobject-detectionObject DetectionRepresentation LearningRetrievalzero-shot-classificationZero-Shot Learning

Similar Papers 제목 키워드 기반

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models

2024-11-28 · Jisheng Bai, Haohe Liu, Mou Wang, Dongyuan Shi 외

With the emergence of audio-language models, constructing large-scale paired audio-language datasets has become essential yet challenging for model development, primarily due to the time-intensive and labour-heavy demand…

Audio captioningAudio to Text RetrievalCaption GenerationRetrieval+1

WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research

2023-03-30 · Xinhao Mei, Chutong Meng, Haohe Liu, Qiuqiang Kong 외

The advancement of audio-language (AL) multimodal learning tasks has been significant in recent years. However, researchers face challenges due to the costly and time-consuming collection process of existing audio-langua…

Audio captioningEvent DetectionLanguage ModellingLarge Language Model+4

A Multi-Purpose Audio-Visual Corpus for Multi-Modal Persian Speech Recognition: the Arman-AV Dataset

2023-01-21 · Javad Peymanfard, Samin Heydarian, Ali Lashini, Hossein Zeinali 외

In recent years, significant progress has been made in automatic lip reading. But these methods require large-scale datasets that do not exist for many low-resource languages. In this paper, we have presented a new multi…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Lip Reading+4

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion

2025-06-01 · Shunian Chen, Xinyuan Xie, Zheshu Chen, Liyan Zhao 외

High-quality, large-scale audio captioning is crucial for advancing audio understanding, yet current automated methods often generate captions that lack fine-grained detail and contextual accuracy, primarily due to their…

Audio captioningCaption GenerationInstruction FollowingLarge Language Model

Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models

2023-01-30 · Rongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren 외

Large-scale multimodal generative modeling has created milestones in text-to-image and text-to-video generation. Its application to audio still lags behind for two main reasons: the lack of large-scale datasets with high…

Audio GenerationText-to-Video GenerationVideo Generation