paper-with-me

Papers

Moodifier: MLLM-Enhanced Emotion-Driven Image Editing

2025-07-18 · Jiarong Ye, Sharon X. Huang arxiv

Bridging emotions and visual content for emotion-driven image editing holds great potential in creative industries, yet precise manipulation remains challenging due to the abstract nature of emotions and their varied manifestations across different contexts. We tackle this challenge with an integrated approach consisting of three complementary components. First, we introduce MoodArchive, an 8M+ image dataset with detailed hierarchical emotional annotations generated by LLaVA and partially validated by human evaluators. Second, we develop MoodifyCLIP, a vision-language model fine-tuned on MoodArchive to translate abstract emotions into specific visual attributes. Third, we propose Moodifier, a training-free editing model leveraging MoodifyCLIP and multimodal large language models (MLLMs) to enable precise emotional transformations while preserving content integrity. Our system works across diverse domains such as character expressions, fashion design, jewelry, and home décor, enabling creators to quickly visualize emotional variations while preserving identity and structure. Extensive experimental evaluations show that Moodifier outperforms existing methods in both emotional accuracy and content preservation, providing contextually appropriate edits. By linking abstract emotions to concrete visual changes, our solution unlocks new possibilities for emotional content creation in real-world applications. We will release the MoodArchive dataset, MoodifyCLIP model, and make the Moodifier code and demo publicly available upon acceptance.

📄 PDF Abstract BibTeX arXiv:2507.14024

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models

2025-05-16 · Bohao Xing, Xin Liu, Guoying Zhao, Chengyu Liu 외

Emotion understanding is a critical yet challenging task. Recent advances in Multimodal Large Language Models (MLLMs) have significantly enhanced their capabilities in this area. However, MLLMs often suffer from hallucin…

Hallucination

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement

2026-07-23 · Daiqing Wu, Dongbao Yang, Jiashu Yao, Hongrui Zhang 외 arxiv

Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward Artificial General Intelligence (AGI). However, despite the rapid pr…

Emotional IntelligenceEmotion Interpretation

MELLM: Exploring LLM-Powered Micro-Expression Understanding Enhanced by Subtle Motion Perception

2025-05-11 · Zhengye Zhang, Sirui Zhao, Shifeng Liu, Shukang Yin 외

Micro-expressions (MEs) are crucial psychological responses with significant potential for affective computing. However, current automatic micro-expression recognition (MER) research primarily focuses on discrete emotion…

Emotion ClassificationLarge Language ModelMicro Expression RecognitionMicro-Expression Recognition+2

MultiEmo-Bench: Multi-label Visual Emotion Analysis for Multi-modal Large Language Models

2026-05-14 · Tianwei Chen, Takuya Furusawa, Yuki Hirakawa, Ryotaro Shimizu 외 arxiv

This paper introduces a multi-label visual emotion analysis benchmark dataset for comprehensively evaluating the ability of multimodal large language models (MLLMs) to predict the emotions evoked by images. Recent user s…

Multimodal Video Emotion Recognition with Reliable Reasoning Priors

2025-07-29 · Zhepeng Wang, Yingjian Zhu, Guanghao Dong, Hongzhu Yi 외 arxiv

This study investigates the integration of trustworthy prior reasoning knowledge from MLLMs into multimodal emotion recognition. We employ Gemini to generate fine-grained, modality-separable reasoning traces, which are i…

Multimodal Emotion RecognitionVideo Emotion RecognitionContrastive Learning