EmoCAM: Toward Understanding What Drives CNN-based Emotion Recognition
Convolutional Neural Networks are particularly suited for image analysis tasks, such as Image Classification, Object Recognition or Image Segmentation. Like all Artificial Neural Networks, however, they are "black box" models, and suffer from poor explainability. This work is concerned with the specific downstream task of Emotion Recognition from images, and proposes a framework that combines CAM-based techniques with Object Detection on a corpus level to better understand on which image cues a particular model, in our case EmoNet, relies to assign a specific emotion to an image. We demonstrate that the model mostly focuses on human characteristics, but also explore the pronounced effect of specific image modifications.
Code (0)
등록된 구현이 없습니다.
Tasks
Emotion Recognitionimage-classificationImage ClassificationImage SegmentationObjectobject-detectionObject DetectionObject RecognitionSemantic SegmentationSimilar Papers 제목 키워드 기반
Multi-Task Learning and Adapted Knowledge Models for Emotion-Cause Extraction
Detecting what emotions are expressed in text is a well-studied problem in natural language processing. However, research on finer grained emotion analysis such as what causes an emotion is still in its infancy. We prese…
Common Sense ReasoningEmotion Cause ExtractionEmotion ClassificationEmotion Recognition+1Emotion Recognition in Context
Understanding what a person is experiencing from her frame of reference is essential in our everyday life. For this reason, one can think that machines with this type of ability would interact better with people. However…
Emotion RecognitionEmotion Recognition in ContextA comparative study of emotion recognition methods using facial expressions
Understanding the facial expressions of our interlocutor is important to enrich the communication and to give it a depth that goes beyond the explicitly expressed. In fact, studying one's facial expression gives insight …
Emotion RecognitionFacial Emotion RecognitionWhat Drives Test-Time Adaptation for CLIP? A Controlled Empirical Study from an Update Perspective
Vision-Language Models (VLMs) such as CLIP have become a standard backbone for open-vocabulary recognition, yet their zero-shot predictions remain vulnerable to distribution shifts encountered at deployment. Test-Time Ad…
Test-time AdaptationGraph Neural Networks for Image Understanding Based on Multiple Cues: Group Emotion Recognition and Event Recognition as Use Cases
A graph neural network (GNN) for image understanding based on multiple cues is proposed in this paper. Compared to traditional feature and decision fusion approaches that neglect the fact that features can interact and e…
Emotion RecognitionGraph Neural Network