paper-with-me

홈 › Papers

Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding

2025-09-18 · Zhu Li, Xiyuan Gao, Yuqing Zhang, Shekhar Nayak, Matt Coler arxiv

Sarcasm detection remains a challenge in natural language understanding, as sarcastic intent often relies on subtle cross-modal cues spanning text, speech, and vision. While prior work has primarily focused on textual or visual-textual sarcasm, comprehensive audio-visual-textual sarcasm understanding remains underexplored. In this paper, we systematically evaluate large language models (LLMs) and multimodal LLMs for sarcasm detection on English (MUStARD++) and Chinese (MCSD 1.0) in zero-shot, few-shot, and LoRA fine-tuning settings. In addition to direct classification, we explore models as feature encoders, integrating their representations through a collaborative gating fusion module. Experimental results show that audio-based models achieve the strongest unimodal performance, while text-audio and audio-vision combinations outperform unimodal and trimodal models. Furthermore, MLLMs such as Qwen-Omni show competitive zero-shot and fine-tuned performance. Our findings highlight the potential of MLLMs for cross-lingual, audio-visual-textual sarcasm understanding.

📄 PDF Abstract BibTeX arXiv:2509.15476

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language UnderstandingSarcasm Detection

Similar Papers 제목 키워드 기반

Seeing Sarcasm Through Different Eyes: Analyzing Multimodal Sarcasm Perception in Large Vision-Language Models

2025-03-15 · Junjie Chen, Xuyang Liu, Subin Huang, Linfeng Zhang 외

With the advent of large vision-language models (LVLMs) demonstrating increasingly human-like abilities, a pivotal question emerges: do different LVLMs interpret multimodal sarcasm differently, and can a single model gra…

Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects

2025-09-04 · Xiyuan Gao, Shekhar Nayak, Matt Coler arxiv

Sarcasm, a common feature of human communication, poses challenges in interpersonal interactions and human-machine interactions. Linguistic research has highlighted the importance of prosodic cues, such as variations in …

Sarcasm Detection

Evaluating Open-Source Vision-Language Models for Multimodal Sarcasm Detection

2025-10-13 · Saroj Basnet, Shafkat Farabi, Tharindu Ranasinghe, Diptesh Kanoji 외 arxiv

Recent advances in open-source vision-language models (VLMs) offer new opportunities for understanding complex and subjective multimodal phenomena such as sarcasm. In this work, we evaluate seven state-of-the-art VLMs - …

Sarcasm Detection

TextMI: Textualize Multimodal Information for Integrating Non-verbal Cues in Pre-trained Language Models

2023-03-27 · Md Kamrul Hasan, Md Saiful Islam, Sangwu Lee, Wasifur Rahman 외

Pre-trained large language models have recently achieved ground-breaking performance in a wide variety of language understanding tasks. However, the same model can not be applied to multimodal behavior understanding task…

Humor DetectionMultimodal Sentiment AnalysisSarcasm DetectionSentiment Analysis

"When Words Fail, Emojis Prevail": Generating Sarcastic Utterances with Emoji Using Valence Reversal and Semantic Incongruity

2023-05-06 · Faria Binte Kader, Nafisa Hossain Nujat, Tasmia Binte Sogir, Mohsinul Kabir 외

Sarcasm is a form of figurative language that serves as a humorous tool for mockery and ridicule. We present a novel architecture for sarcasm generation with emoji from a non-sarcastic input sentence in English. We divid…

General KnowledgeSentence