EmotionCaps: Enhancing Audio Captioning Through Emotion-Augmented Data Generation
Recent progress in audio-language modeling, such as automated audio captioning, has benefited from training on synthetic data generated with the aid of large-language models. However, such approaches for environmental sound captioning have primarily focused on audio event tags and have not explored leveraging emotional information that may be present in recordings. In this work, we explore the benefit of generating emotion-augmented synthetic audio caption data by instructing ChatGPT with additional acoustic information in the form of estimated soundscape emotion. To do so, we introduce EmotionCaps, an audio captioning dataset comprised of approximately 120,000 audio clips with paired synthetic descriptions enriched with soundscape emotion recognition (SER) information. We hypothesize that this additional information will result in higher-quality captions that match the emotional tone of the audio recording, which will, in turn, improve the performance of captioning models trained with this data. We test this hypothesis through both objective and subjective evaluation, comparing models trained with the EmotionCaps dataset to multiple baseline models. Our findings challenge current approaches to captioning and suggest new directions for developing and assessing captioning models.
Code (0)
등록된 구현이 없습니다.
Tasks
Audio captioningEmotion RecognitionLanguage ModelingLanguage ModellingSimilar Papers 제목 키워드 기반
Bridging Paintings and Music -- Exploring Emotion based Music Generation through Paintings
Rapid advancements in artificial intelligence have significantly enhanced generative tasks involving music and images, employing both unimodal and multimodal approaches. This research develops a model capable of generati…
FADImage CaptioningMusic Generationtext similarityADIFF: Explaining audio difference using natural language
Understanding and explaining differences between audio recordings is crucial for fields like audio forensics, quality assessment, and audio generation. This involves identifying and describing audio events, acoustic scen…
AudioCapsAudio captioningAudio GenerationLanguage Modeling+2Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
Current vision-guided audio captioning systems frequently fail to address audiovisual misalignment in real-world scenarios, such as dubbed content or off-screen sounds. To bridge this critical gap, we present an entropy-…
AudioCapsAudio captioningUncertainty QuantificationEnhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding
Automated audio captioning (AAC) is an audio-to-text task to describe audio contents in natural language. Recently, the advancements in large language models (LLMs), with improvements in training approaches for audio enc…
Audio captioningDecoderSPECTRUM: Semantic Processing and Emotion-informed video-Captioning Through Retrieval and Understanding Modalities
Capturing a video's meaning and critical concepts by analyzing the subtle details is a fundamental yet challenging task in video captioning. Identifying the dominant emotional tone in a video significantly enhances the p…
AttributeDescriptiveRetrievalText Retrieval+2