paper-with-me

홈 › Papers

EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation

2026-08-06 · Bingyuan Wang, Baistan Zhyldyzbekov, Kunyu Feng, Zeyu Wang arxiv

Emotion shapes how viewers interpret a scene, yet existing video generators entangle global atmosphere, affect-bearing semantic cues, and temporal progression within a single text condition. We present EmoWorld, a framework that decouples these factors within a frozen flow-matching video diffusion transformer (Video DiT). A one-time preparation stage extracts layer-specific affect directions and a reusable cue library from geometry-preserving neutral and emotion-edited panoramas. At inference, Visual Atmosphere Steering (VAS) injects atmosphere directions into hidden states, Semantic Affective Steering (SAS) isolates a separately scalable prompt residual for semantic cues, and Temporal Affective Steering (TAS) interpolates endpoint residual fields across denoising and video time. On Wan2.2, VAS improves target-emotion alignment by 19% while reducing a temporal-fluctuation proxy by 48%; SAS improves target-emotion alignment by 37% and increases detected affect-bearing cues by 36%; and TAS improves transition monotonicity by 15% over the strongest baseline. EmoWorld is evaluated across 27 emotion categories in text-to-video and image-to-video settings, demonstrates portability across multiple Video-DiT backbones, and supports camera-conditioned composition without updating generator parameters.

📄 PDF Abstract BibTeX arXiv:2608.06231

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction

2026-08-06 · Jiahao Huang, Zheng Lian, Jingyi Zhang, Zhide Chen 외 hf

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However, prevailing research predominantly focuses on task-specific specialization, often neglecting inter-tas…

Reinforcement LearningEmotional IntelligenceMultimodal ReasoningMulti-Task Learning

Controllable Affective Generation via Latent Vector Steering

2026-08-26 · Xixian Yong, Siyuan Chang, Yingying Zhang, Xian Wu 외 arxiv

Large Language Models (LLMs) often produce emotionally flattened responses after alignment, limiting their effectiveness in affect-sensitive applications. In this paper, we propose EmoVec, a lightweight framework for con…

Continuous Control

EmoCtrl: Controllable Emotional Image Content Generation

2025-12-27 · Jingyuan Yang, Weibin Luo, Hui Huang arxiv

An image conveys meaning through both its visual content and emotional tone, jointly shaping human perception. We introduce Controllable Emotional Image Content Generation (C-EICG), which aims to generate images that rem…

CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering

2026-02-03 · Siyi Wang, Shihong Tan, Siyi Liu, Hong Jia 외 arxiv

Emotional expression in human speech is nuanced and compositional, often involving multiple, sometimes conflicting, affective cues that may diverge from linguistic content. In contrast, most expressive text-to-speech sys…

Steering Emotional Dynamics for Art Therapy: Controllable Narrative Script Generation through Hierarchically Guided LLM Agents

2026-06-15 · Suqing Wang, Qinghai Miao, Chao Guo, Yisheng Lv arxiv

Art therapy plays a vital role in emotional healing, in which narrative creation acts as the primary vehicle for emotional expression. Given the inherently dynamic nature of emotions during healing, narratives with finel…

Trajectory PlanningScene Generation