paper-with-me

Papers

Complementary Fusion of Multi-Features and Multi-Modalities in Sentiment Analysis

2019-04-17 · Feiyang Chen, Ziqian Luo, Yanyan Xu, Dengfeng Ke

Sentiment analysis, mostly based on text, has been rapidly developing in the last decade and has attracted widespread attention in both academia and industry. However, the information in the real world usually comes from multiple modalities, such as audio and text. Therefore, in this paper, based on audio and text, we consider the task of multimodal sentiment analysis and propose a novel fusion strategy including both multi-feature fusion and multi-modality fusion to improve the accuracy of audio-text sentiment analysis. We call it the DFF-ATMF (Deep Feature Fusion - Audio and Text Modality Fusion) model, which consists of two parallel branches, the audio modality based branch and the text modality based branch. Its core mechanisms are the fusion of multiple feature vectors and multiple modality attention. Experiments on the CMU-MOSI dataset and the recently released CMU-MOSEI dataset, both collected from YouTube for sentiment analysis, show the very competitive results of our DFF-ATMF model. Furthermore, by virtue of attention weight distribution heatmaps, we also demonstrate the deep features learned by using DFF-ATMF are complementary to each other and robust. Surprisingly, DFF-ATMF also achieves new state-of-the-art results on the IEMOCAP dataset, indicating that the proposed fusion strategy also has a good generalization ability for multimodal emotion recognition.

📄 PDF Abstract BibTeX arXiv:1904.08138

Code (3)

robertjkeck2/EmoNet
robertjkeck2/EmoTe
robertjkeck2/pwdl

Tasks

Emotion RecognitionMultimodal Emotion RecognitionMultimodal Sentiment AnalysisSentiment Analysis

Similar Papers 제목 키워드 기반

Inconsistency-Aware Cross-Attention for Audio-Visual Fusion in Dimensional Emotion Recognition

2024-05-21 · G Rajasekhar, Jahangir Alam

Leveraging complementary relationships across modalities has recently drawn a lot of attention in multimodal emotion recognition. Most of the existing approaches explored cross-attention to capture the complementary rela…

Emotion RecognitionMultimodal Emotion Recognition

MMIF-AMIN: Adaptive Loss-Driven Multi-Scale Invertible Dense Network for Multimodal Medical Image Fusion

2025-08-12 · Tao Luo, Weihua Xu arxiv

Multimodal medical image fusion (MMIF) aims to integrate images from different modalities to produce a comprehensive image that enhances medical diagnosis by accurately depicting organ structures, tissue textures, and me…

Medical Diagnosis

A Joint Cross-Attention Model for Audio-Visual Fusion in Dimensional Emotion Recognition

2022-03-28 · R. Gnana Praveen, Wheidima Carneiro de Melo, Nasib Ullah, Haseeb Aslam 외

Multimodal emotion recognition has recently gained much attention since it can leverage diverse and complementary relationships over multiple modalities (e.g., audio, visual, biosignals, etc.), and can provide some robus…

Emotion RecognitionMultimodal Emotion Recognition

CrossFuse: A Novel Cross Attention Mechanism based Infrared and Visible Image Fusion Approach

2024-06-15 · Hui Li, Xiao-Jun Wu

Multimodal visual information fusion aims to integrate the multi-sensor data into a single image which contains more complementary information and less redundant features. However the complementary information is hard to…

DecoderInfrared And Visible Image Fusion

BiCo-Fusion: Bidirectional Complementary LiDAR-Camera Fusion for Semantic- and Spatial-Aware 3D Object Detection

2024-06-27 · Yang song, Lin Wang

3D object detection is an important task that has been widely applied in autonomous driving. To perform this task, a new trend is to fuse multi-modal inputs, i.e., LiDAR and camera. Under such a trend, recent methods fus…

3D Object DetectionAutonomous DrivingImage Enhancementobject-detection+1