paper-with-me

홈 › Papers

Knowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion Recognition

2025-05-26 · CVPR 2025 1 · Wen Yin, Yong Wang, Guiduo Duan, Dongyang Zhang, Xin Hu, Yuan-Fang Li, Tao He

Visual Emotion Recognition (VER) is a critical yet challenging task aimed at inferring emotional states of individuals based on visual cues. However, existing works focus on single domains, e.g., realistic images or stickers, limiting VER models' cross-domain generalizability. To fill this gap, we introduce an Unsupervised Cross-Domain Visual Emotion Recognition (UCDVER) task, which aims to generalize visual emotion recognition from the source domain (e.g., realistic images) to the low-resource target domain (e.g., stickers) in an unsupervised manner. Compared to the conventional unsupervised domain adaptation problems, UCDVER presents two key challenges: a significant emotional expression variability and an affective distribution shift. To mitigate these issues, we propose the Knowledge-aligned Counterfactual-enhancement Diffusion Perception (KCDP) framework. Specifically, KCDP leverages a VLM to align emotional representations in a shared knowledge space and guides diffusion models for improved visual affective perception. Furthermore, a Counterfactual-Enhanced Language-image Emotional Alignment (CLIEA) method generates high-quality pseudo-labels for the target domain. Extensive experiments demonstrate that our model surpasses SOTA models in both perceptibility and generalization, e.g., gaining 12% improvements over the SOTA VER model TGCA-PVT. The project page is at https://yinwen2019.github.io/ucdver.

📄 PDF Abstract BibTeX arXiv:2505.19694

Code (0)

등록된 구현이 없습니다.

Tasks

counterfactualDomain AdaptationEmotion RecognitionUnsupervised Domain Adaptation

Methods 이 논문이 사용한 방법론

Focus 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Diffusion Counterfactual Generation with Semantic Abduction

2025-06-09 · Rajat Rasal, Avinash Kori, Fabio De Sousa Ribeiro, Tian Xia 외

Counterfactual image generation presents significant challenges, including preserving identity, maintaining perceptual quality, and ensuring faithfulness to an underlying causal model. While existing auto-encoding framew…

counterfactualCounterfactual ReasoningImage GenerationRepresentation Learning

CFPO: Counterfactual Policy Optimization for Multimodal Reasoning

2026-06-22 · Zhangyuan Yu, Wanran Sun, Guangjing Yang, Xiaohu Wu 외 arxiv

Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in multimodal reasoning. However, prevailing reinforcement learning (RL) paradigms lack explicit counterfactual enhancement and causal learni…

Reinforcement LearningMultimodal Reasoning

Latent Diffusion Counterfactual Explanations

2023-10-10 · Karim Farid, Simon Schrodi, Max Argus, Thomas Brox

Counterfactual explanations have emerged as a promising method for elucidating the behavior of opaque black-box models. Recently, several works leveraged pixel-space diffusion models for counterfactual generation. To han…

counterfactual

Imagining Alternatives: Towards High-Resolution 3D Counterfactual Medical Image Generation via Language Guidance

2025-09-07 · Mohamed Mohamed, Brennan Nichyporuk, Douglas L. Arnold, Tal Arbel arxiv

Vision-language models have demonstrated impressive capabilities in generating 2D images under various conditions; however, the success of these models is largely enabled by extensive, readily available pretrained founda…

Medical Image Generation

Low-light Image Enhancement via CLIP-Fourier Guided Wavelet Diffusion

2024-01-08 · Minglong Xue, Jinhong He, Wenhai Wang, Mingliang Zhou

Low-light image enhancement techniques have significantly progressed, but unstable image quality recovery and unsatisfactory visual perception are still significant challenges. To solve these problems, we propose a novel…

Image EnhancementLow-Light Image Enhancement