paper-with-me

홈 › Papers

Emotion and Intent Joint Understanding in Multimodal Conversation: A Benchmarking Dataset

2024-07-03 · Rui Liu, Haolin Zuo, Zheng Lian, Xiaofen Xing, Björn W. Schuller, Haizhou Li

Emotion and Intent Joint Understanding in Multimodal Conversation (MC-EIU) aims to decode the semantic information manifested in a multimodal conversational history, while inferring the emotions and intents simultaneously for the current utterance. MC-EIU is enabling technology for many human-computer interfaces. However, there is a lack of available datasets in terms of annotation, modality, language diversity, and accessibility. In this work, we propose an MC-EIU dataset, which features 7 emotion categories, 9 intent categories, 3 modalities, i.e., textual, acoustic, and visual content, and two languages, i.e., English and Mandarin. Furthermore, it is completely open-source for free access. To our knowledge, MC-EIU is the first comprehensive and rich emotion and intent joint understanding dataset for multimodal conversation. Together with the release of the dataset, we also develop an Emotion and Intent Interaction (EI$^2$) network as a reference system by modeling the deep correlation between emotion and intent in the multimodal conversation. With comparative experiments and ablation studies, we demonstrate the effectiveness of the proposed EI$^2$ method on the MC-EIU dataset. The dataset and codes will be made available at: https://github.com/MC-EIU/MC-EIU.

📄 PDF Abstract BibTeX arXiv:2407.02751

Code (1)

mc-eiu/mc-eiu 공식 구현 pytorch

Tasks

BenchmarkingDiversity

Similar Papers 제목 키워드 기반

A$^2$-LLM: An End-to-end Conversational Audio Avatar Large Language Model

2026-02-04 · Xiaolin Hu, Hang Yuan, Xinzhu Sang, Binbin Yan 외 arxiv

Developing expressive and responsive conversational digital humans is a cornerstone of next-generation human-computer interaction. While large language models (LLMs) have significantly enhanced dialogue capabilities, mos…

Uncertain Multimodal Intention and Emotion Understanding in the Wild

2025-01-01 · CVPR 2025 1 · Qu Yang, Qinghongya Shi, Tongxin Wang, Mang Ye

Understanding intention and emotion from social media poses unique challenges due to the inherent uncertainty in multimodal data, where posts often contain incomplete or missing modalities. While this uncertainty ref…

Affective Multimodal Agents with Proactive Knowledge Grounding for Emotionally Aligned Marketing Dialogue

2025-11-21 · Lin Yu, Xiaofei Han, Yifei Kang, Chiung-Yi Tseng 외 arxiv

Recent advances in large language models (LLMs) have enabled fluent dialogue systems, but most remain reactive and struggle in emotionally rich, goal-oriented settings such as marketing conversations. To address this lim…

Multimodal Emotion-Cause Pair Extraction in Conversations

2021-10-15 · Fanfan Wang, Zixiang Ding, Rui Xia, Zhaoyu Li 외

Emotion cause analysis has received considerable attention in recent years. Previous studies primarily focused on emotion cause extraction from texts in news articles or microblogs. It is also interesting to discover emo…

ArticlesEmotion Cause ExtractionEmotion-Cause Pair ExtractionEmotion Recognition+1

AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis

2026-07-17 · Zhenqi Jia, Yuan Zhao, Aruukhan, Rui Liu 외 arxiv

Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent interactions. Existing CSS methods struggle to render authentic human emotions…

Speech Synthesis