paper-with-me

홈 › Papers

Using Scene and Semantic Features for Multi-modal Emotion Recognition

2023-08-01 · Zhifeng Wang, Ramesh Sankaranarayana

Automatic emotion recognition is a hot topic with a wide range of applications. Much work has been done in the area of automatic emotion recognition in recent years. The focus has been mainly on using the characteristics of a person such as speech, facial expression and pose for this purpose. However, the processing of scene and semantic features for emotion recognition has had limited exploration. In this paper, we propose to use combined scene and semantic features, along with personal features, for multi-modal emotion recognition. Scene features will describe the environment or context in which the target person is operating. The semantic feature can include objects that are present in the environment, as well as their attributes and relationships with the target person. In addition, we use a modified EmbraceNet to extract features from the images, which is trained to learn both the body and pose features simultaneously. By fusing both body and pose features, the EmbraceNet can improve the accuracy and robustness of the model, particularly when dealing with partially missing data. This is because having both body and pose features provides a more complete representation of the subject in the images, which can help the model to make more accurate predictions even when some parts of body are missing. We demonstrate the efficiency of our method on the benchmark EMOTIC dataset. We report an average precision of 40.39\% across the 26 emotion categories, which is a 5\% improvement over previous approaches.

📄 PDF Abstract BibTeX arXiv:2308.00228

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion Recognition

Methods 이 논문이 사용한 방법론

EmbraceNet 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

SEER: Semantic Enhancement and Emotional Reasoning Network for Multimodal Fake News Detection

2025-07-17 · Peican Zhu, Yubo Jing, Le Cheng, Bin Chen 외 arxiv

Previous studies on multimodal fake news detection mainly focus on the alignment and integration of cross-modal features, as well as the application of text-image consistency. However, they overlook the semantic enhancem…

Fake News Detection

Towards Accurate Emotion-Attributed Video Captioning via Fine-grained Emotion-Cause Pair Extraction

2026-06-07 · Weidong Chen, Cheng Ye, Zhendong Mao, Liping Wang 외 arxiv

Emotional Video Captioning (EVC) is a challenging task that aims to generate factually accurate and emotionally rich descriptions for videos. Existing EVC methods leverage holistic visual features to mine global emotiona…

Emotion-Cause Pair ExtractionVideo Captioning

Adversarial Representation with Intra-Modal and Inter-Modal Graph Contrastive Learning for Multimodal Emotion Recognition

2023-12-28 · Yuntao Shou, Tao Meng, Wei Ai, Nan Yin 외

With the release of increasing open-source emotion recognition datasets on social media platforms and the rapid development of computing resources, multimodal emotion recognition tasks (MER) have begun to receive widespr…

Contrastive LearningEmotion RecognitionGraph Representation LearningMultimodal Emotion Recognition+1

DER-GCN: Dialogue and Event Relation-Aware Graph Convolutional Neural Network for Multimodal Dialogue Emotion Recognition

2023-12-17 · Wei Ai, Yuntao Shou, Tao Meng, Nan Yin 외

With the continuous development of deep learning (DL), the task of multimodal dialogue emotion recognition (MDER) has recently received extensive research attention, which is also an essential branch of DL. The MDER aims…

Contrastive LearningEmotion RecognitionMultimodal Emotion RecognitionRepresentation Learning

Affect-Prototype Guided Fusion for Open-Vocabulary Incomplete Multi-modal Emotion Recognition

2026-09-15 · Yichi Zhang, Shenyue Wang, Jing Luo, Chunyang Yu 외 arxiv

Open-vocabulary multimodal emotion recognition (OV-MER) aims to generate open natural-language emotion labels from multimodal affective cues. In real-world scenarios, however, complete and synchronized modal data are dif…

Multimodal Emotion Recognition