paper-with-me

Papers

UniEmoX: Cross-modal Semantic-Guided Large-Scale Pretraining for Universal Scene Emotion Perception

2024-09-27 · Chuang Chen, Xiao Sun, Zhi Liu

Visual emotion analysis holds significant research value in both computer vision and psychology. However, existing methods for visual emotion analysis suffer from limited generalizability due to the ambiguity of emotion perception and the diversity of data scenarios. To tackle this issue, we introduce UniEmoX, a cross-modal semantic-guided large-scale pretraining framework. Inspired by psychological research emphasizing the inseparability of the emotional exploration process from the interaction between individuals and their environment, UniEmoX integrates scene-centric and person-centric low-level image spatial structural information, aiming to derive more nuanced and discriminative emotional representations. By exploiting the similarity between paired and unpaired image-text samples, UniEmoX distills rich semantic knowledge from the CLIP model to enhance emotional embedding representations more effectively. To the best of our knowledge, this is the first large-scale pretraining framework that integrates psychological theories with contemporary contrastive learning and masked image modeling techniques for emotion analysis across diverse scenarios. Additionally, we develop a visual emotional dataset titled Emo8. Emo8 samples cover a range of domains, including cartoon, natural, realistic, science fiction and advertising cover styles, covering nearly all common emotional scenes. Comprehensive experiments conducted on six benchmark datasets across two downstream tasks validate the effectiveness of UniEmoX. The source code is available at https://github.com/chincharles/u-emo.

📄 PDF Abstract BibTeX arXiv:2409.18877

Code (1)

chincharles/u-emo 공식 구현 pytorch

Tasks

Contrastive LearningEmotion Recognition

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

SGMA: Semantic-Guided Modality-Aware Segmentation for Remote Sensing with Incomplete Multimodal Data

2026-03-03 · Lekang Wen, Liang Liao, Jing Xiao, Mi Wang arxiv

Multimodal semantic segmentation integrates complementary information from diverse sensors for remote sensing Earth observation. However, practical systems often encounter missing modalities due to sensor failures or inc…

Semantic SegmentationContrastive Learning

UniGeo: A Multi-modal Large Language Model for Text-Guided Cross-View Geo-Localization

2026-08-27 · Jiahao Wen, Hang Yu, Zhedong Zheng arxiv

Text-guided drone geo-localization aims to identify a target region in a large-scale image gallery from a natural-language description. Existing methods mainly formulate this task as direct matching between an open-ended…

Wnet: Audio-Guided Video Object Segmentation via Wavelet-Based Cross-Modal Denoising Networks

2022-01-01 · CVPR 2022 1 · Wenwen Pan, Haonan Shi, Zhou Zhao, Jieming Zhu 외

Audio-Guided video semantic segmentation is a challenging problem in visual analysis and editing, which automatically separates foreground objects from background in a video sequence according to the referring audio …

DecoderDenoisingSegmentationSemantic Segmentation+2

Uncertainty-Resilient Multimodal Learning via Consistency-Guided Cross-Modal Transfer

2025-11-18 · Hyo-Jeong Jang arxiv

Multimodal learning systems often face substantial uncertainty due to noisy data, low-quality labels, and heterogeneous modality characteristics. These issues become especially critical in human-computer interaction sett…

Representation Learning

Large Language Model Guided Progressive Feature Alignment for Multimodal UAV Object Detection

2025-03-10 · Wentao Wu, Chenglong Li, Xiao Wang, Bin Luo 외

Existing multimodal UAV object detection methods often overlook the impact of semantic gaps between modalities, which makes it difficult to achieve accurate semantic and spatial alignments, limiting detection performance…

Language ModelingLanguage ModellingLarge Language ModelObject+2