paper-with-me

Papers

EmotionCaps: Enhancing Audio Captioning Through Emotion-Augmented Data Generation

2024-10-15 · Mithun Manivannan, Vignesh Nethrapalli, Mark Cartwright

Recent progress in audio-language modeling, such as automated audio captioning, has benefited from training on synthetic data generated with the aid of large-language models. However, such approaches for environmental sound captioning have primarily focused on audio event tags and have not explored leveraging emotional information that may be present in recordings. In this work, we explore the benefit of generating emotion-augmented synthetic audio caption data by instructing ChatGPT with additional acoustic information in the form of estimated soundscape emotion. To do so, we introduce EmotionCaps, an audio captioning dataset comprised of approximately 120,000 audio clips with paired synthetic descriptions enriched with soundscape emotion recognition (SER) information. We hypothesize that this additional information will result in higher-quality captions that match the emotional tone of the audio recording, which will, in turn, improve the performance of captioning models trained with this data. We test this hypothesis through both objective and subjective evaluation, comparing models trained with the EmotionCaps dataset to multiple baseline models. Our findings challenge current approaches to captioning and suggest new directions for developing and assessing captioning models.

📄 PDF Abstract BibTeX arXiv:2410.12028

Code (0)

등록된 구현이 없습니다.

Tasks

Audio captioningEmotion RecognitionLanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Bridging Paintings and Music -- Exploring Emotion based Music Generation through Paintings

2024-09-12 · Tanisha Hisariya, huan zhang, Jinhua Liang

Rapid advancements in artificial intelligence have significantly enhanced generative tasks involving music and images, employing both unimodal and multimodal approaches. This research develops a model capable of generati…

FADImage CaptioningMusic Generationtext similarity

ADIFF: Explaining audio difference using natural language

2025-02-06 · Soham Deshmukh, Shuo Han, Rita Singh, Bhiksha Raj

Understanding and explaining differences between audio recordings is crucial for fields like audio forensics, quality assessment, and audio generation. This involves identifying and describing audio events, acoustic scen…

AudioCapsAudio captioningAudio GenerationLanguage Modeling+2

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning

2025-05-28 · Le Xu, Chenxing Li, Yong Ren, Yujie Chen 외

Current vision-guided audio captioning systems frequently fail to address audiovisual misalignment in real-world scenarios, such as dubbed content or off-screen sounds. To bridge this critical gap, we present an entropy-…

AudioCapsAudio captioningUncertainty Quantification

Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding

2024-06-19 · Jizhong Liu, Gang Li, Junbo Zhang, Heinrich Dinkel 외

Automated audio captioning (AAC) is an audio-to-text task to describe audio contents in natural language. Recently, the advancements in large language models (LLMs), with improvements in training approaches for audio enc…

Audio captioningDecoder

SPECTRUM: Semantic Processing and Emotion-informed video-Captioning Through Retrieval and Understanding Modalities

2024-11-04 · Ehsan Faghihi, Mohammedreza Zarenejad, Ali-Asghar Beheshti Shirazi

Capturing a video's meaning and critical concepts by analyzing the subtle details is a fundamental yet challenging task in video captioning. Identifying the dominant emotional tone in a video significantly enhances the p…

AttributeDescriptiveRetrievalText Retrieval+2