paper-with-me

Papers

Why We Feel: Breaking Boundaries in Emotional Reasoning with Multimodal Large Language Models

2025-04-10 · Yuxiang Lin, Jingdong Sun, Zhi-Qi Cheng, Jue Wang, Haomin Liang, Zebang Cheng, Yifei Dong, Jun-Yan He, Xiaojiang Peng, Xian-Sheng Hua

Most existing emotion analysis emphasizes which emotion arises (e.g., happy, sad, angry) but neglects the deeper why. We propose Emotion Interpretation (EI), focusing on causal factors-whether explicit (e.g., observable objects, interpersonal interactions) or implicit (e.g., cultural context, off-screen events)-that drive emotional responses. Unlike traditional emotion recognition, EI tasks require reasoning about triggers instead of mere labeling. To facilitate EI research, we present EIBench, a large-scale benchmark encompassing 1,615 basic EI samples and 50 complex EI samples featuring multifaceted emotions. Each instance demands rationale-based explanations rather than straightforward categorization. We further propose a Coarse-to-Fine Self-Ask (CFSA) annotation pipeline, which guides Vision-Language Models (VLLMs) through iterative question-answer rounds to yield high-quality labels at scale. Extensive evaluations on open-source and proprietary large language models under four experimental settings reveal consistent performance gaps-especially for more intricate scenarios-underscoring EI's potential to enrich empathetic, context-aware AI applications. Our benchmark and methods are publicly available at: https://github.com/Lum1104/EIBench, offering a foundation for advanced multimodal causal analysis and next-generation affective computing.

📄 PDF Abstract BibTeX arXiv:2504.07521

Code (2)

lum1104/eibench 공식 구현 pytorch
Lum1104/MER-Factory

Tasks

Emotion InterpretationEmotion Recognition

Similar Papers 제목 키워드 기반

StimuVAR: Spatiotemporal Stimuli-aware Video Affective Reasoning with Multimodal Large Language Models

2024-08-31 · Yuxiang Guo, Faizan Siddiqui, Yang Zhao, Rama Chellappa 외

Predicting and reasoning how a video would make a human feel is crucial for developing socially intelligent systems. Although Multimodal Large Language Models (MLLMs) have shown impressive video understanding capabilitie…

Video Understanding

Empathetic Response Generation through Graph-based Multi-hop Reasoning on Emotional Causality

2021-10-09 · Jiashuo Wang, Wenjie Li, Peiqin Lin, Feiteng Mu

Empathetic response generation aims to comprehend the user emotion and then respond to it appropriately. Most existing works merely focus on what the emotion is and ignore how the emotion is evoked, thus weakening the ca…

Empathetic Response GenerationResponse Generation

Emotion Recognition With Temporarily Localized 'Emotional Events' in Naturalistic Context

2022-10-25 · Mohammad Asif, Sudhakar Mishra, Majithia Tejas Vinodbhai, Uma Shanker Tiwary

Emotion recognition using EEG signals is an emerging area of research due to its broad applicability in BCI. Emotional feelings are hard to stimulate in the lab. Emotions do not last long, yet they need enough context to…

EEGElectroencephalogram (EEG)Emotion Recognition

Socratis: Are large multimodal models emotionally aware?

2023-08-31 · Katherine Deng, Arijit Ray, Reuben Tan, Saadia Gabriel 외

Existing emotion prediction benchmarks contain coarse emotion labels which do not consider the diversity of emotions that an image and text can elicit in humans due to various reasons. Learning diverse reactions to multi…

Articles

FEEL: A Framework for Evaluating Emotional Support Capability with Large Language Models

2024-03-23 · Huaiwen Zhang, Yu Chen, Ming Wang, Shi Feng

Emotional Support Conversation (ESC) is a typical dialogue that can effectively assist the user in mitigating emotional pressures. However, owing to the inherent subjectivity involved in analyzing emotions, current non-a…

Ensemble Learning