paper-with-me

Papers

Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective Model

2024-09-01 · Fuqiang Niu, Zebang Cheng, Xianghua Fu, Xiaojiang Peng, Genan Dai, Yin Chen, Hu Huang, BoWen Zhang

Stance detection, which aims to identify public opinion towards specific targets using social media data, is an important yet challenging task. With the proliferation of diverse multimodal social media content including text, and images multimodal stance detection (MSD) has become a crucial research area. However, existing MSD studies have focused on modeling stance within individual text-image pairs, overlooking the multi-party conversational contexts that naturally occur on social media. This limitation stems from a lack of datasets that authentically capture such conversational scenarios, hindering progress in conversational MSD. To address this, we introduce a new multimodal multi-turn conversational stance detection dataset (called MmMtCSD). To derive stances from this challenging dataset, we propose a novel multimodal large language model stance detection framework (MLLM-SD), that learns joint stance representations from textual and visual modalities. Experiments on MmMtCSD show state-of-the-art performance of our proposed MLLM-SD approach for multimodal stance detection. We believe that MmMtCSD will contribute to advancing real-world applications of stance detection research.

📄 PDF Abstract BibTeX arXiv:2409.00597

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language ModelStance Detection

Similar Papers 제목 키워드 기반

C-MTCSD: A Chinese Multi-Turn Conversational Stance Detection Dataset

2025-04-14 · Fuqiang Niu, Yi Yang, Xianghua Fu, Genan Dai 외

Stance detection has become an essential tool for analyzing public discussions on social media. Current methods face significant challenges, particularly in Chinese language processing and multi-turn conversational analy…

Stance Detection

MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations

2024-03-16 · Hanlei Zhang, Xin Wang, Hua Xu, Qianrui Zhou 외

Multimodal intent recognition poses significant challenges, requiring the incorporation of non-verbal modalities from real-world contexts to enhance the comprehension of human intentions. Existing benchmark datasets are …

Intent RecognitionMultimodal Intent Recognition

From Videos to Conversations: Egocentric Instructions for Task Assistance

2026-02-01 · Lavisha Aggarwal, Vikas Bahirwani, Andrea Colaco arxiv

Many everyday tasks, ranging from appliance repair and cooking to car maintenance, require expert knowledge, particularly for complex, multi-step procedures. Despite growing interest in AI agents for augmented reality (A…

Duplex Conversation: Towards Human-like Interaction in Spoken Dialogue Systems

2022-05-30 · Ting-En Lin, Yuchuan Wu, Fei Huang, Luo Si 외

In this paper, we present Duplex Conversation, a multi-turn, multimodal spoken dialogue system that enables telephone-based agents to interact with customers like a human. We use the concept of full-duplex in telecommuni…

Data AugmentationSpoken Dialogue Systems

Evaluating Large Language Models Abilities for Addressee, Turn-change, and Next Speaker Prediction in Meetings

2026-06-16 · Ryo Fukuda, Takatomo Kano, Siddhant Arora, Marc Delcroix 외 arxiv

We investigate turn-taking in multimodal multi-party conversations using large language models (LLMs). We construct an evaluation framework for three tasks: addressee detection, turn-change prediction, and next speaker p…