paper-with-me

홈 › Papers

M2FNet: Multi-modal Fusion Network for Emotion Recognition in Conversation

2022-06-05 · Vishal Chudasama, Purbayan Kar, Ashish Gudmalwar, Nirmesh Shah, Pankaj Wasnik, Naoyuki Onoe

Emotion Recognition in Conversations (ERC) is crucial in developing sympathetic human-machine interaction. In conversational videos, emotion can be present in multiple modalities, i.e., audio, video, and transcript. However, due to the inherent characteristics of these modalities, multi-modal ERC has always been considered a challenging undertaking. Existing ERC research focuses mainly on using text information in a discussion, ignoring the other two modalities. We anticipate that emotion recognition accuracy can be improved by employing a multi-modal approach. Thus, in this study, we propose a Multi-modal Fusion Network (M2FNet) that extracts emotion-relevant features from visual, audio, and text modality. It employs a multi-head attention-based fusion mechanism to combine emotion-rich latent representations of the input data. We introduce a new feature extractor to extract latent features from the audio and visual modality. The proposed feature extractor is trained with a novel adaptive margin-based triplet loss function to learn emotion-relevant features from the audio and visual data. In the domain of ERC, the existing methods perform well on one benchmark dataset but not on others. Our results show that the proposed M2FNet architecture outperforms all other methods in terms of weighted average F1 score on well-known MELD and IEMOCAP datasets and sets a new state-of-the-art performance in ERC.

📄 PDF Abstract BibTeX arXiv:2206.02187

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionEmotion Recognition in ConversationTriplet

Methods 이 논문이 사용한 방법론

Triplet Loss The goal of Triplet loss, in the context of Siamese Networks, is to maximize the joint probability among all score-pairs i.e. the product of all probabilities. By using its…

Similar Papers 제목 키워드 기반

BPFNet: A Unified Framework for Bimodal Palmprint Alignment and Fusion

2021-10-04 · Zhaoqun Li, Xu Liang, Dandan Fan, Jinxing Li 외

Bimodal palmprint recognition leverages palmprint and palm vein images simultaneously,which achieves high accuracy by multi-model information fusion and has strong anti-falsification property. In the recognition pipeline…

Keypoint DetectionTranslation

Multi-level Attention Fusion Network for Audio-visual Event Recognition

2021-06-12 · Mathilde Brousmiche, Jean Rouat, Stéphane Dupont

Event classification is inherently sequential and multimodal. Therefore, deep neural models need to dynamically focus on the most relevant time window and/or modality of a video. In this study, we propose the Multi-level…

ERIT Lightweight Multimodal Dataset for Elderly Emotion Recognition and Multimodal Fusion Evaluation

2024-07-25 · Rita Frieske, Bertrand E. Shi

ERIT is a novel multimodal dataset designed to facilitate research in a lightweight multimodal fusion. It contains text and image data collected from videos of elderly individuals reacting to various situations, as well …

Emotion Recognition

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition

2024-12-07 · Feng Li, Jiusong Luo, Wanjun Xia

Speech emotion recognition (SER) remains a challenging yet crucial task due to the inherent complexity and diversity of human emotions. To address this problem, researchers attempt to fuse information from other modaliti…

DiversityEmotion RecognitionRepresentation LearningSpeech Emotion Recognition

Audio-visual Speaker Recognition with a Cross-modal Discriminative Network

2020-08-10

Audio-visual speaker recognition is one of the tasks in the recent 2019 NIST speaker recognition evaluation (SRE). Studies in neuroscience and computer science all point to the fact that vision and auditory neural signal…

Speaker Recognition