paper-with-me

Papers

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization

2026-07-23 · Lihuang Fang, Yuchen Zou, kebin Jin, Jinghui Qin arxiv

Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognition (MER) tasks and lifted MER to a new level that is complex emotion understanding with advanced video understanding abilities and natural language description. However, existing MLLM-based methods often use a fixed prompt to perceive the emotions, ignoring the dynamicity and complexity of the emotion source in the multimodal inputs. To address these issues, we propose a novel Reinforcement Learning-based Dynamic Agent Specialization framework (\textbf{EmoAgent-R1}) to optimize the emotion recognition, reasoning, and generalization abilities of an MLLM with dynamic agent specialization based on reinforcement learning. Specifically, we first adopt a cold start strategy to endow an MLLM with preliminary emotion recognition, reasoning, and agent routing ability by training with synthetic answer-conditioned chain-of-thought data and agent routing data. Then, we further train the MLLM with reinforcement learning to perceive emotions in a two-step agentic workflow with agent selection and agent specialization. To effectively train EmoAgent-R1, we propose a novel Progressive Group-Relative Policy Optimization (P-GRPO) to combine group-based relative advantages with a PMI-inspired progressive token-level modulation to transform sparse rewards into fine-grained learning signals, mitigating the coarse-grained uniform credit assignment issue in GRPO. Extensive experiments on MER benchmarks demonstrate the superiority of our EmoAgent-R1 in stronger emotion reasoning performance and improved optimization stability.

📄 PDF Abstract BibTeX arXiv:2607.21013

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Emotion RecognitionReinforcement Learning

Similar Papers 제목 키워드 기반

EmoAgent: A Multi-Agent Framework for Diverse Affective Image Manipulation

2025-03-14 · Qi Mao, Haobo Hu, Yujie He, Difei Gao 외

Affective Image Manipulation (AIM) aims to alter visual elements within an image to evoke specific emotional responses from viewers. However, existing AIM approaches rely on rigid \emph{one-to-one} mappings between emoti…

Decision MakingImage Manipulation

The Emotional Baby Is Truly Deadly: Does your Multimodal Large Reasoning Model Have Emotional Flattery towards Humans?

2025-08-06 · Yuan Xun, Xiaojun Jia, Xinwei Liu, Hua Zhang arxiv

We observe that MLRMs oriented toward human-centric service are highly susceptible to user emotional cues during the deep-thinking stage, often overriding safety protocols or built-in safety checks under high emotional i…

EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safety

2025-04-13 · Jiahao Qiu, Yinghui He, Xinzhe Juan, Yimin Wang 외

The rise of LLM-driven AI characters raises safety concerns, particularly for vulnerable human users with psychological disorders. To address these risks, we propose EmoAgent, a multi-agent AI framework designed to evalu…

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models

2026-02-27 · Yiyang Fang, Wenke Huang, Pei Fu, Yihao Yang 외 arxiv

Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual reasoning and understanding tasks but still struggle to capture the complexity and subjectivity of human emotions. Existing approaches bas…

Emotional IntelligenceReinforcement LearningVisual Reasoning

MM-DFN: Multimodal Dynamic Fusion Network for Emotion Recognition in Conversations

2022-03-04 · Dou Hu, Xiaolong Hou, Lingwei Wei, Lianxin Jiang 외

Emotion Recognition in Conversations (ERC) has considerable prospects for developing empathetic machines. For multimodal ERC, it is vital to understand context and fuse modality information in conversations. Recent graph…

Emotion RecognitionEmotion Recognition in Conversation