paper-with-me

홈 › Papers

VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models

2025-11-04 · Zhicheng Zhang, Weicheng Wang, Yongjie Zhu, Wenyu Qin, Pengfei Wan, Di Zhang, Jufeng Yang arxiv

Understanding and predicting emotion from videos has gathered significant attention in recent studies, driven by advancements in video large language models (VideoLLMs). While advanced methods have made progress in video emotion analysis, the intrinsic nature of emotions poses significant challenges. Emotions are characterized by dynamic and cues-dependent properties, making it difficult to understand complex and evolving emotional states with reasonable rationale. To tackle these challenges, we propose a novel affective cues-guided reasoning framework that unifies fundamental attribute perception, expression analysis, and high-level emotional understanding in a stage-wise manner. At the core of our approach is a family of video emotion foundation models (VidEmo), specifically designed for emotion reasoning and instruction-following. These models undergo a two-stage tuning process: first, curriculum emotion learning for injecting emotion knowledge, followed by affective-tree reinforcement learning for emotion reasoning. Moreover, we establish a foundational data infrastructure and introduce a emotion-centric fine-grained dataset (Emo-CFG) consisting of 2.1M diverse instruction-based samples. Emo-CFG includes explainable emotional question-answering, fine-grained captions, and associated rationales, providing essential resources for advancing emotion understanding tasks. Experimental results demonstrate that our approach achieves competitive performance, setting a new milestone across 15 face perception tasks.

📄 PDF Abstract BibTeX arXiv:2511.02712

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Benchmarking Dynamic Affective Reasoning: A Viewer-Centric Video Emotion Dataset

2026-07-11 · Zhiyan Zhang, Peipei Song, Jinpeng Hu, Jingyang Jia 외 arxiv

Video emotion analysis is typically framed as a static classification problem, treating each clip as an independent labeled unit. However, such a formulation overlooks a key psychological fact: emotions change as a resul…

Emotion Classification

AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis

2026-04-07 · Dong She, Xianrong Yao, Liqun Chen, Jinghe Yu 외 arxiv

Vision-Language Models (VLMs) have demonstrated strong capabilities in perception, yet holistic Affective Image Content Analysis (AICA), which integrates perception, reasoning, and generation into a unified framework, re…

KEVER^2: Knowledge-Enhanced Visual Emotion Reasoning and Retrieval

2025-05-30 · Fanhang Man, Xiaoyue Chen, Huandong Wang, Baining Zhao 외

Understanding what emotions images evoke in their viewers is a foundational goal in human-centric visual computing. While recent advances in vision-language models (VLMs) have shown promise for visual emotion analysis (V…

Emotion RecognitionRetrieval

Dual-Model Prediction of Affective Engagement and Vocal Attractiveness from Speaker Expressiveness in Video Learning

2026-03-19 · Hung-Yue Suen, Kuo-En Hung, Fan-Hsun Tseng arxiv

This paper outlines a machine learning-enabled speaker-centric Emotion AI approach capable of predicting audience-affective engagement and vocal attractiveness in asynchronous video-based learning, relying solely on spea…

Affective Visual Dialog: A Large-Scale Benchmark for Emotional Reasoning Based on Visually Grounded Conversations

2023-08-30 · Kilichbek Haydarov, Xiaoqian Shen, Avinash Madasu, Mahmoud Salem 외

We introduce Affective Visual Dialog, an emotion explanation and reasoning task as a testbed for research on understanding the formation of emotions in visually grounded conversations. The task involves three skills: (1)…

Explanation GenerationQuestion AnsweringVisual Dialog