paper-with-me

홈 › Papers

MERRY: Semantically Decoupled Evaluation of Multimodal Emotional and Role Consistencies of Role-Playing Agents

2026-02-24 · Zhenyu Wang, Xiaofen Xing, Yirong Chen, Xiangmin Xu arxiv

Multimodal Role-Playing Agents (MRPAs) are attracting increasing attention due to their ability to deliver more immersive multimodal emotional interactions. However, existing studies still rely on pure textual benchmarks to evaluate the text responses of MRPAs, while delegating the assessment of their multimodal expressions solely to modality-synthesis metrics. This evaluation paradigm, on the one hand, entangles semantic assessment with modality generation, leading to ambiguous error attribution, and on the other hand remains constrained by the heavy reliance on human judgment. To this end, we propose MERRY, a semantically decoupled evaluation framework for assessing Multimodal Emotional and Role consistencies of Role-playing agents. This framework introduce five refined metrics for EC and three for RC. Notably, we transform the traditional subjective scoring approach into a novel bidirectional-evidence-finding task, significantly improving the human agreement of LLM-as-Judge evaluations. Based on MERRY, we conduct extensive evaluations. Our empirical results primarily reveal that: (1) Training on synthetic datasets tends to reduce emotional consistency, whereas training on real-world datasets improves it; (2) Existing models suffer from emotional templatization and simplification, exhibiting positive-bias and performance bottleneck in fine-grained negative emotions; (3) Simple prompting method strengthens the weak models but constrains the strong ones, while simple fine-tuning method suffers from poor role generalization. Codes and dataset are available.

📄 PDF Abstract BibTeX arXiv:2602.21941

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Emotion Collider: Dual Hyperbolic Mirror Manifolds for Sentiment Recovery via Anti Emotion Reflection

2026-02-18 · Rong Fu, Ziming Wang, Shuo Yin, Haiyun Wei 외 arxiv

Emotional expression underpins natural communication and effective human-computer interaction. We present Emotion Collider (EC-Net), a hyperbolic hypergraph framework for multimodal emotion and sentiment modeling. EC-Net…

Contrastive Learning

DRKF: Decoupled Representations with Knowledge Fusion for Multimodal Emotion Recognition

2025-08-03 · Peiyuan Jiang, Yao Liu, Qiao Liu, Zongshun Zhang 외 arxiv

Multimodal emotion recognition (MER) aims to identify emotional states by integrating and analyzing information from multiple modalities. However, inherent modality heterogeneity and inconsistencies in emotional cues rem…

Multimodal Emotion RecognitionRepresentation LearningEmotion Classification

RoseMerry: A Baseline Message-level Sentiment Classification System

2015-06-01 · SEMEVAL 2015 6 · Huizhi Liang, Richard Fothergill, Timothy Baldwin
ClassificationGeneral ClassificationRepresentation LearningSentiment Analysis+1

Bridging the Emotional Semantic Gap via Multimodal Relevance Estimation

2023-02-03 · Chuan Zhang, Daoxin Zhang, Ruixiu Zhang, Jiawei Li 외

Human beings have rich ways of emotional expressions, including facial action, voice, and natural languages. Due to the diversity and complexity of different individuals, the emotions expressed by various modalities may …

Contrastive Learning

SPECTRUM: Semantic Processing and Emotion-informed video-Captioning Through Retrieval and Understanding Modalities

2024-11-04 · Ehsan Faghihi, Mohammedreza Zarenejad, Ali-Asghar Beheshti Shirazi

Capturing a video's meaning and critical concepts by analyzing the subtle details is a fundamental yet challenging task in video captioning. Identifying the dominant emotional tone in a video significantly enhances the p…

AttributeDescriptiveRetrievalText Retrieval+2