paper-with-me

Papers

Enriching Multimodal Sentiment Analysis through Textual Emotional Descriptions of Visual-Audio Content

2024-12-12 · Sheng Wu, Xiaobao Wang, Longbiao Wang, Dongxiao He, Jianwu Dang

Multimodal Sentiment Analysis (MSA) stands as a critical research frontier, seeking to comprehensively unravel human emotions by amalgamating text, audio, and visual data. Yet, discerning subtle emotional nuances within audio and video expressions poses a formidable challenge, particularly when emotional polarities across various segments appear similar. In this paper, our objective is to spotlight emotion-relevant attributes of audio and visual modalities to facilitate multimodal fusion in the context of nuanced emotional shifts in visual-audio scenarios. To this end, we introduce DEVA, a progressive fusion framework founded on textual sentiment descriptions aimed at accentuating emotional features of visual-audio content. DEVA employs an Emotional Description Generator (EDG) to transmute raw audio and visual data into textualized sentiment descriptions, thereby amplifying their emotional characteristics. These descriptions are then integrated with the source data to yield richer, enhanced features. Furthermore, DEVA incorporates the Text-guided Progressive Fusion Module (TPF), leveraging varying levels of text as a core modality guide. This module progressively fuses visual-audio minor modalities to alleviate disparities between text and visual-audio modalities. Experimental results on widely used sentiment analysis benchmark datasets, including MOSI, MOSEI, and CH-SIMS, underscore significant enhancements compared to state-of-the-art models. Moreover, fine-grained emotion experiments corroborate the robust sensitivity of DEVA to subtle emotional variations.

📄 PDF Abstract BibTeX arXiv:2412.10460

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Sentiment AnalysisSentiment Analysis

Similar Papers 제목 키워드 기반

Picturized and Recited with Dialects: A Multimodal Chinese Representation Framework for Sentiment Analysis of Classical Chinese Poetry

2025-05-19 · Xiaocong Du, Haoyu Pei, Haipeng Zhang

Classical Chinese poetry is a vital and enduring part of Chinese literature, conveying profound emotional resonance. Existing studies analyze sentiment based on textual meanings, overlooking the unique rhythmic and visua…

Representation LearningSentenceSentiment Analysis

Leveraging Textual-Cues for Enhancing Multimodal Sentiment Analysis by Object Recognition

2026-01-30 · Sumana Biswas, Karen Young, Josephine Griffith arxiv

Multimodal sentiment analysis, which includes both image and text data, presents several challenges due to the dissimilarities in the modalities of text and image, the ambiguity of sentiment, and the complexities of cont…

Multimodal Sentiment AnalysisObject Recognition

A Novel Context-Aware Multimodal Framework for Persian Sentiment Analysis

2021-03-03 · Kia Dashtipour, Mandar Gogate, Erik Cambria, Amir Hussain

Most recent works on sentiment analysis have exploited the text modality. However, millions of hours of video recordings posted on social media platforms everyday hold vital unstructured information that can be exploited…

Multimodal Sentiment AnalysisPersian Sentiment AnalysisSentiment Analysis

PSA-MF: Personality-Sentiment Aligned Multi-Level Fusion for Multimodal Sentiment Analysis

2025-12-01 · Heng Xie, Kang Zhu, Zhengqi Wen, Jianhua Tao 외 arxiv

Multimodal sentiment analysis (MSA) is a research field that recognizes human sentiments by combining textual, visual, and audio modalities. The main challenge lies in integrating sentiment-related information from diffe…

Multimodal Sentiment Analysis

Sentiment Word Aware Multimodal Refinement for Multimodal Sentiment Analysis with ASR Errors

2022-03-01 · Findings (ACL) 2022 5 · Yang Wu, Yanyan Zhao, Hao Yang, Song Chen 외

Multimodal sentiment analysis has attracted increasing attention and lots of models have been proposed. However, the performance of the state-of-the-art models decreases sharply when they are deployed in the real world. …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Multimodal Sentiment AnalysisPosition+4