Recognizing Emotions in Video Using Multimodal DNN Feature Fusion
We present our system description of input-level multimodal fusion of audio, video, and text for recognition of emotions and their intensities for the 2018 First Grand Challenge on Computational Modeling of Human Multimodal Language. Our proposed approach is based on input-level feature fusion with sequence learning from Bidirectional Long-Short Term Memory (BLSTM) deep neural networks (DNNs). We show that our fusion approach outperforms unimodal predictors. Our system performs 6-way simultaneous classification and regression, allowing for overlapping emotion labels in a video segment. This leads to an overall binary accuracy of 90{\%}, overall 4-class accuracy of 89.2{\%} and an overall mean-absolute-error (MAE) of 0.12. Our work shows that an early fusion technique can effectively predict the presence of multi-label emotions as well as their coarse-grained intensities. The presented multimodal approach creates a simple and robust baseline on this new Grand Challenge dataset. Furthermore, we provide a detailed analysis of emotion intensity distributions as output from our DNN, as well as a related discussion concerning the inherent difficulty of this task.
Code (0)
등록된 구현이 없습니다.
Tasks
Emotion RecognitionMachine TranslationSimilar Papers 제목 키워드 기반
Hierarchical Audio-Visual Information Fusion with Multi-label Joint Decoding for MER 2023
In this paper, we propose a novel framework for recognizing both discrete and dimensional emotions. In our framework, deep features extracted from foundation models are used as robust acoustic and visual representations …
Emotion ClassificationEmotion RecognitionMultimodal Emotion RecognitionregressionAnalyzing the Influence of Dataset Composition for Emotion Recognition
Recognizing emotions from text in multimodal architectures has yielded promising results, surpassing video and audio modalities under certain circumstances. However, the method by which multimodal data is collected can b…
Emotion RecognitionMultimodal Emotion RecognitionUnimodal-driven Distillation in Multimodal Emotion Recognition with Dynamic Fusion
Multimodal Emotion Recognition in Conversations (MERC) identifies emotional states across text, audio and video, which is essential for intelligent dialogue systems and opinion analysis. Existing methods emphasize hetero…
Emotion RecognitionKnowledge DistillationMixture-of-ExpertsMultimodal Emotion RecognitionTextualized and Feature-based Models for Compound Multimodal Emotion Recognition in the Wild
Systems for multimodal emotion recognition (ER) are commonly trained to extract features from different modalities (e.g., visual, audio, and textual) that are combined to predict individual basic emotions. However, compo…
Emotion RecognitionMultimodal Emotion RecognitionConversational Memory Network for Emotion Recognition in Dyadic Dialogue Videos
Emotion recognition in conversations is crucial for the development of empathetic machines. Present methods mostly ignore the role of inter-speaker dependency relations while classifying emotions in conversations. In thi…
Emotion RecognitionEmotion Recognition in Conversation