paper-with-me

홈 › Papers

EmoGen: Emotional Image Content Generation with Text-to-Image Diffusion Models

2024-01-09 · CVPR 2024 1 · Jingyuan Yang, Jiawei Feng, Hui Huang

Recent years have witnessed remarkable progress in image generation task, where users can create visually astonishing images with high-quality. However, existing text-to-image diffusion models are proficient in generating concrete concepts (dogs) but encounter challenges with more abstract ones (emotions). Several efforts have been made to modify image emotions with color and style adjustments, facing limitations in effectively conveying emotions with fixed image contents. In this work, we introduce Emotional Image Content Generation (EICG), a new task to generate semantic-clear and emotion-faithful images given emotion categories. Specifically, we propose an emotion space and construct a mapping network to align it with the powerful Contrastive Language-Image Pre-training (CLIP) space, providing a concrete interpretation of abstract emotions. Attribute loss and emotion confidence are further proposed to ensure the semantic diversity and emotion fidelity of the generated images. Our method outperforms the state-of-the-art text-to-image approaches both quantitatively and qualitatively, where we derive three custom metrics, i.e., emotion accuracy, semantic clarity and semantic diversity. In addition to generation, our method can help emotion understanding and inspire emotional art design.

📄 PDF Abstract BibTeX arXiv:2401.04608

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeDiversityImage Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content Generation

2025-08-05 · Kaishen Yuan, Yuting Zhang, Shang Gao, Yijie Zhu 외 arxiv

Emotional Image Content Generation (EICG) aims to generate semantically clear and emotionally faithful images based on given emotion categories, with broad application prospects. While recent text-to-image diffusion mode…

EmoGene: Audio-Driven Emotional 3D Talking-Head Generation

2024-10-07 · Wenqing Wang, Yun Fu

Audio-driven talking-head generation is a crucial and useful technology for virtual human interaction and film-making. While recent advances have focused on improving image fidelity and lip synchronization, generating ac…

NeRFTalking Head Generation

EmoGen: Eliminating Subjective Bias in Emotional Music Generation

2023-07-03 · Chenfei Kang, Peiling Lu, Botao Yu, Xu Tan 외

Music is used to convey emotions, and thus generating emotional music is important in automatic music generation. Previous work on emotional music generation directly uses annotated emotion labels as control signals, whi…

AttributeClusteringMusic GenerationSelf-Supervised Learning

MemoGen: Can Past Experience Improve Future Text-to-Image Generation?

2026-06-02 · Wenshuo Chen, Kuimou Yu, Bowen Tian, Jianfei Song 외 arxiv

Modern text-to-image models have achieved strong visual synthesis, yet remain unreliable when prompts require implicit visual constraints, relational reasoning, or external knowledge. Existing retrieval-augmented and age…

Text-to-Image GenerationRelational ReasoningContinual Learning

FineMoGen: Fine-Grained Spatio-Temporal Motion Generation and Editing

2023-12-22 · NeurIPS 2023 11 · Mingyuan Zhang, Huirong Li, Zhongang Cai, Jiawei Ren 외

Text-driven motion generation has achieved substantial progress with the emergence of diffusion models. However, existing methods still struggle to generate complex motion sequences that correspond to fine-grained descri…

Mixture-of-ExpertsMotion GenerationMotion Synthesis