paper-with-me

Papers

KEVER^2: Knowledge-Enhanced Visual Emotion Reasoning and Retrieval

2025-05-30 · Fanhang Man, Xiaoyue Chen, Huandong Wang, Baining Zhao, Han Li, Xinlei Chen, Yong Li

Understanding what emotions images evoke in their viewers is a foundational goal in human-centric visual computing. While recent advances in vision-language models (VLMs) have shown promise for visual emotion analysis (VEA), several key challenges remain unresolved. Emotional cues in images are often abstract, overlapping, and entangled, making them difficult to model and interpret. Moreover, VLMs struggle to align these complex visual patterns with emotional semantics due to limited supervision and sparse emotional grounding. Finally, existing approaches lack structured affective knowledge to resolve ambiguity and ensure consistent emotional reasoning across diverse visual domains. To address these limitations, we propose \textbf{K-EVER\textsuperscript{2}}, a knowledge-enhanced framework for emotion reasoning and retrieval. Our approach introduces a semantically structured formulation of visual emotion cues and integrates external affective knowledge through multimodal alignment. Without relying on handcrafted labels or direct emotion supervision, K-EVER\textsuperscript{2} achieves robust and interpretable emotion predictions across heterogeneous image types. We validate our framework on three representative benchmarks, Emotion6, EmoSet, and M-Disaster, covering social media imagery, human-centric scenes, and disaster contexts. K-EVER\textsuperscript{2} consistently outperforms strong CNN and VLM baselines, achieving up to a \textbf{19\% accuracy gain} for specific emotions and a \textbf{12.3\% average accuracy gain} across all emotion categories. Our results demonstrate a scalable and generalizable solution for advancing emotional understanding of visual content.

📄 PDF Abstract BibTeX arXiv:2505.24342

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionRetrieval

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations

2025-06-01 · Parul Gupta, Shreya Ghosh, Tom Gedeon, Thanh-Toan Do 외

The rapid advancement of GenAI technology over the past few years has significantly contributed towards highly realistic deepfake content generation. Despite ongoing efforts, the research community still lacks a large-sc…

DeepFake DetectionFace SwappingHuman-Object Interaction Detection

SOLVER: Scene-Object Interrelated Visual Emotion Reasoning Network

2021-10-24 · Jingyuan Yang, Xinbo Gao, Leida Li, Xiumei Wang 외

Visual Emotion Analysis (VEA) aims at finding out how people feel emotionally towards different visual stimuli, which has attracted great attention recently with the prevalence of sharing images on social networks. Since…

Emotion RecognitionObject

Neutral Utterances are Also Causes: Enhancing Conversational Causal Emotion Entailment with Social Commonsense Knowledge

2022-05-02 · Jiangnan Li, Fandong Meng, Zheng Lin, Rui Liu 외

Conversational Causal Emotion Entailment aims to detect causal utterances for a non-neutral targeted utterance from a conversation. In this work, we build conversations as graphs to overcome implicit contextual modelling…

Causal Emotion Entailment

Multimodal Video Emotion Recognition with Reliable Reasoning Priors

2025-07-29 · Zhepeng Wang, Yingjian Zhu, Guanghao Dong, Hongzhu Yi 외 arxiv

This study investigates the integration of trustworthy prior reasoning knowledge from MLLMs into multimodal emotion recognition. We employ Gemini to generate fine-grained, modality-separable reasoning traces, which are i…

Multimodal Emotion RecognitionVideo Emotion RecognitionContrastive Learning

Knowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion Recognition

2025-05-26 · CVPR 2025 1 · Wen Yin, Yong Wang, Guiduo Duan, Dongyang Zhang 외

Visual Emotion Recognition (VER) is a critical yet challenging task aimed at inferring emotional states of individuals based on visual cues. However, existing works focus on single domains, e.g., realistic images or stic…

counterfactualDomain AdaptationEmotion RecognitionUnsupervised Domain Adaptation