paper-with-me

홈 › Papers

CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content Generation

2025-08-05 · Kaishen Yuan, Yuting Zhang, Shang Gao, Yijie Zhu, Wenshuo Chen, Yutao Yue arxiv

Emotional Image Content Generation (EICG) aims to generate semantically clear and emotionally faithful images based on given emotion categories, with broad application prospects. While recent text-to-image diffusion models excel at generating concrete concepts, they struggle with the complexity of abstract emotions. There have also emerged methods specifically designed for EICG, but they excessively rely on word-level attribute labels for guidance, which suffer from semantic incoherence, ambiguity, and limited scalability. To address these challenges, we propose CoEmoGen, a novel pipeline notable for its semantic coherence and high scalability. Specifically, leveraging multimodal large language models (MLLMs), we construct high-quality captions focused on emotion-triggering content for context-rich semantic guidance. Furthermore, inspired by psychological insights, we design a Hierarchical Low-Rank Adaptation (HiLoRA) module to cohesively model both polarity-shared low-level features and emotion-specific high-level semantics. Extensive experiments demonstrate CoEmoGen's superiority in emotional faithfulness and semantic coherence from quantitative, qualitative, and user study perspectives. To intuitively showcase scalability, we curate EmoArt, a large-scale dataset of emotionally evocative artistic images, providing endless inspiration for emotion-driven artistic creation. The dataset and code are available at https://github.com/yuankaishen2001/CoEmoGen.

📄 PDF Abstract BibTeX arXiv:2508.03535

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation

2025-09-02 · Zikai Huang, Yihan Zhou, Xuemiao Xu, Cheng Xu 외 arxiv

Singing-driven 3D head animation is a challenging yet promising task with applications in virtual avatars, entertainment, and education. Unlike speech, singing involves richer emotional nuance, dynamic prosody, and lyric…

KEVER^2: Knowledge-Enhanced Visual Emotion Reasoning and Retrieval

2025-05-30 · Fanhang Man, Xiaoyue Chen, Huandong Wang, Baining Zhao 외

Understanding what emotions images evoke in their viewers is a foundational goal in human-centric visual computing. While recent advances in vision-language models (VLMs) have shown promise for visual emotion analysis (V…

Emotion RecognitionRetrieval

EmoSEM: Segment and Explain Emotion Stimuli in Visual Art

2025-04-20 · Jing Zhang, Dan Guo, Zhangbin Li, Meng Wang

This paper focuses on a key challenge in visual art understanding: given an art image, the model pinpoints pixel regions that trigger a specific human emotion, and generates linguistic explanations for the emotional arou…

Emotion InterpretationSegmentation

Multiview Contextual Commonsense Inference: A New Dataset and Task

2022-10-06 · Siqi Shen, Deepanway Ghosal, Navonil Majumder, Henry Lim 외

Contextual commonsense inference is the task of generating various types of explanations around the events in a dyadic dialogue, including cause, motivation, emotional reaction, and others. Producing a coherent and non-t…

Multiview Contextual Commonsense Inference

OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory

2025-12-08 · Zhaochong An, Menglin Jia, Haonan Qiu, Zijian Zhou 외 arxiv

Storytelling in real-world videos often unfolds through multiple shots -- discontinuous yet semantically connected clips that together convey a coherent narrative. However, existing multi-shot video generation (MSV) meth…

Video Generation