paper-with-me

홈 › Papers

EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation

2025-11-14 · Zongyang Qiu, Bingyuan Wang, Xingbei Chen, Yingqing He, Zeyu Wang arxiv

Emotion plays a pivotal role in video-based expression, but existing video generation systems predominantly focus on low-level visual metrics while neglecting affective dimensions. Although emotion analysis has made progress in the visual domain, the video community lacks dedicated resources to bridge emotion understanding with generative tasks, particularly for stylized and non-realistic contexts. To address this gap, we introduce EmoVid, the first multimodal, emotion-annotated video dataset specifically designed for creative media, which includes cartoon animations, movie clips, and animated stickers. Each video is annotated with emotion labels, visual attributes (brightness, colorfulness, hue), and text captions. Through systematic analysis, we uncover spatial and temporal patterns linking visual features to emotional perceptions across diverse video forms. Building on these insights, we develop an emotion-conditioned video generation technique by fine-tuning the Wan2.1 model. The results show a significant improvement in both quantitative metrics and the visual quality of generated videos for text-to-video and image-to-video tasks. EmoVid establishes a new benchmark for affective video computing. Our work not only offers valuable insights into visual emotion analysis in artistically styled videos, but also provides practical methods for enhancing emotional expression in video generation.

📄 PDF Abstract BibTeX arXiv:2511.11002

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation

2025-01-03 · CVPR 2025 1 · Huaize Liu, Wenzhang Sun, Donglin Di, Shibo Sun 외

The generation of talking avatars has achieved significant advancements in precise audio synchronization. However, crafting lifelike talking head videos requires capturing a broad spectrum of emotions and subtle facial e…

Portrait Animation

Emotion-Qwen: Training Hybrid Experts for Unified Emotion and General Vision-Language Understanding

2025-05-10 · Dawei Huang, Qing Li, Chuan Yan, Zebang Cheng 외

Emotion understanding in videos aims to accurately recognize and interpret individuals' emotional states by integrating contextual, visual, textual, and auditory cues. While Large Multimodal Models (LMMs) have demonstrat…

DescriptiveEmotion RecognitionMixture-of-Experts

ERIT Lightweight Multimodal Dataset for Elderly Emotion Recognition and Multimodal Fusion Evaluation

2024-07-25 · Rita Frieske, Bertrand E. Shi

ERIT is a novel multimodal dataset designed to facilitate research in a lightweight multimodal fusion. It contains text and image data collected from videos of elderly individuals reacting to various situations, as well …

Emotion Recognition

ICON: Interactive Conversational Memory Network for Multimodal Emotion Detection

2018-10-01 · EMNLP 2018 10 · Devamanyu Hazarika, Soujanya Poria, Rada Mihalcea, Erik Cambria 외

Emotion recognition in conversations is crucial for building empathetic machines. Present works in this domain do not explicitly consider the inter-personal influences that thrive in the emotional dynamics of dialogues. …

Emotion RecognitionEmotion Recognition in ConversationGeneral ClassificationMultimodal Emotion Recognition+1

Multimodal Emotion-Cause Pair Extraction in Conversations

2021-10-15 · Fanfan Wang, Zixiang Ding, Rui Xia, Zhaoyu Li 외

Emotion cause analysis has received considerable attention in recent years. Previous studies primarily focused on emotion cause extraction from texts in news articles or microblogs. It is also interesting to discover emo…

ArticlesEmotion Cause ExtractionEmotion-Cause Pair ExtractionEmotion Recognition+1