paper-with-me

Papers

MicroEmo: Time-Sensitive Multimodal Emotion Recognition with Micro-Expression Dynamics in Video Dialogues

2024-07-23 · Liyun Zhang

Multimodal Large Language Models (MLLMs) have demonstrated remarkable multimodal emotion recognition capabilities, integrating multimodal cues from visual, acoustic, and linguistic contexts in the video to recognize human emotional states. However, existing methods ignore capturing local facial features of temporal dynamics of micro-expressions and do not leverage the contextual dependencies of the utterance-aware temporal segments in the video, thereby limiting their expected effectiveness to a certain extent. In this work, we propose MicroEmo, a time-sensitive MLLM aimed at directing attention to the local facial micro-expression dynamics and the contextual dependencies of utterance-aware video clips. Our model incorporates two key architectural contributions: (1) a global-local attention visual encoder that integrates global frame-level timestamp-bound image features with local facial features of temporal dynamics of micro-expressions; (2) an utterance-aware video Q-Former that captures multi-scale and contextual dependencies by generating visual token sequences for each utterance segment and for the entire video then combining them. Preliminary qualitative experiments demonstrate that in a new Explainable Multimodal Emotion Recognition (EMER) task that exploits multi-modal and multi-faceted clues to predict emotions in an open-vocabulary (OV) manner, MicroEmo demonstrates its effectiveness compared with the latest methods.

📄 PDF Abstract BibTeX arXiv:2407.16552

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionMultimodal Emotion Recognition

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Global-Local Attention 설명 없음

Similar Papers 제목 키워드 기반

Noise-Resistant Multimodal Transformer for Emotion Recognition

2023-05-04 · Yuanyuan Liu, Haoyu Zhang, Yibing Zhan, Zijing Chen 외

Multimodal emotion recognition identifies human emotions from various data modalities like video, text, and audio. However, we found that this task can be easily affected by noisy information that does not contain useful…

Emotion RecognitionMultimodal Emotion Recognition

Implicit Design Choices and Their Impact on Emotion Recognition Model Development and Evaluation

2023-09-06 · Mimansa Jaiswal

Emotion recognition is a complex task due to the inherent subjectivity in both the perception and production of emotions. The subjectivity of emotions poses significant challenges in developing accurate and robust comput…

Data AugmentationEmotion Recognition

MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations

2018-10-05 · ACL 2019 7 · Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik 외

Emotion recognition in conversations is a challenging task that has recently gained popularity due to its potential applications. Until now, however, a large-scale multimodal multi-party emotional conversational database…

Dialogue GenerationEmotion RecognitionEmotion Recognition in Conversation

End-to-End Multimodal Emotion Recognition using Deep Neural Networks

2017-04-27 · Panagiotis Tzirakis, George Trigeorgis, Mihalis A. Nicolaou, Björn Schuller 외

Automatic affect recognition is a challenging task due to the various modalities emotions can be expressed with. Applications can be found in many domains including multimedia retrieval and human computer interaction. In…

Emotion RecognitionMultimodal Emotion RecognitionRetrieval

Learning Alignment for Multimodal Emotion Recognition from Speech

2019-09-06 · Haiyang Xu, HUI ZHANG, Kun Han, Yun Wang 외

Speech emotion recognition is a challenging problem because human convey emotions in subtle and complex ways. For emotion recognition on human speech, one can either extract emotion related features from audio signals or…

Emotion RecognitionMultimodal Emotion RecognitionSpeech Emotion Recognitionspeech-recognition+1