paper-with-me

홈 › Papers

HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs

2025-08-14 · Zheng Qin, Ruobing Zheng, Yabing Wang, Tianqi Li, Yi Yuan, Jingdong Chen, Le Wang arxiv

While Multimodal Large Language Models (MLLMs) show immense promise for achieving truly human-like interactions, progress is hindered by the lack of fine-grained evaluation frameworks for human-centered scenarios, encompassing both the understanding of complex human intentions and the provision of empathetic, context-aware responses. Here we introduce HumanSense, a comprehensive benchmark designed to evaluate the human-centered perception and interaction capabilities of MLLMs, with a particular focus on deep understanding of extended multimodal contexts and the formulation of rational feedback. Our evaluation reveals that leading MLLMs still have considerable room for improvement, particularly for advanced interaction-oriented tasks. Supplementing visual input with audio and text information yields substantial improvements, and Omni-modal models show advantages on these tasks.Furthermore, grounded in the observation that appropriate feedback stems from a contextual analysis of the interlocutor's needs and emotions, we posit that reasoning ability serves as the key to unlocking it. We devise a multi-stage, modality-progressive reinforcement learning approach, resulting in HumanSense-Omni-Reasoning, which substantially enhances performance on higher-level understanding and interactive tasks. Additionally, we observe that successful reasoning processes appear to exhibit consistent thought patterns. By designing corresponding prompts, we also enhance the performance of non-reasoning models in a training-free manner.Project page: \textcolor{brightpink}{https://digital-avatar.github.io/ai/HumanSense/}

📄 PDF Abstract BibTeX arXiv:2508.10576

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

A Multi-Agent Framework with Structured Reasoning and Reflective Refinement for Multimodal Empathetic Response Generation

2026-04-21 · Liping Wang, Cheng Ye, Weidong Chen, Peipei Song 외 arxiv

Multimodal empathetic response generation (MERG) aims to generate emotionally engaging and empathetic responses based on users' multimodal contexts. Existing approaches usually rely on an implicit one-pass generation par…

Empathetic Response Generation

AIVA: An AI-based Virtual Companion for Emotion-aware Interaction

2025-09-03 · Chenxi Li arxiv

Recent advances in Large Language Models (LLMs) have significantly improved natural language understanding and generation, enhancing Human-Computer Interaction (HCI). However, LLMs are limited to unimodal text processing…

Natural Language UnderstandingContrastive LearningPrompt Engineering

E-CORE: Emotion Correlation Enhanced Empathetic Dialogue Generation

2023-11-25 · Fengyi Fu, Lei Zhang, Quan Wang, Zhendong Mao

Achieving empathy is a crucial step toward humanized dialogue systems. Current approaches for empathetic dialogue generation mainly perceive an emotional label to generate an empathetic response conditioned on it, which …

DecoderDialogue GenerationResponse Generation

Empathetic Response Generation via Emotion Cause Transition Graph

2023-02-23 · Yushan Qian, Bo wang, Ting-En Lin, Yinhe Zheng 외

Empathetic dialogue is a human-like behavior that requires the perception of both affective factors (e.g., emotion status) and cognitive factors (e.g., cause of the emotion). Besides concerning emotion status in early wo…

DecoderEmpathetic Response GenerationResponse Generation

STICKERCONV: Generating Multimodal Empathetic Responses from Scratch

2024-01-20 · Yiqun Zhang, Fanheng Kong, Peidong Wang, Shuang Sun 외

Stickers, while widely recognized for enhancing empathetic communication in online interactions, remain underexplored in current empathetic dialogue research, notably due to the challenge of a lack of comprehensive datas…

2kEmpathetic Response GenerationResponse Generation