paper-with-me

Papers

MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm Detection

2025-10-27 · Haochen Zhao, Yuyao Kong, Yongxiu Xu, Gaopeng Gou, Hongbo Xu, Yubin Wang, Haoliang Zhang arxiv

Despite progress in multimodal sarcasm detection, existing datasets and methods predominantly focus on single-image scenarios, overlooking potential semantic and affective relations across multiple images. This leaves a gap in modeling cases where sarcasm is triggered by multi-image cues in real-world settings. To bridge this gap, we introduce MMSD3.0, a new benchmark composed entirely of multi-image samples curated from tweets and Amazon reviews. We further propose the Cross-Image Reasoning Model (CIRM), which performs targeted cross-image sequence modeling to capture latent inter-image connections. In addition, we introduce a relevance-guided, fine-grained cross-modal fusion mechanism based on text-image correspondence to reduce information loss during integration. We establish a comprehensive suite of strong and representative baselines and conduct extensive experiments, showing that MMSD3.0 is an effective and reliable benchmark that better reflects real-world conditions. Moreover, CIRM demonstrates state-of-the-art performance across MMSD, MMSD2.0 and MMSD3.0, validating its effectiveness in both single-image and multi-image scenarios. Dataset and code are publicly available at https://github.com/ZHCMOONWIND/MMSD3.0.

📄 PDF Abstract BibTeX arXiv:2510.23299

Code (0)

등록된 구현이 없습니다.

Tasks

Sarcasm Detection

Similar Papers 제목 키워드 기반

MMSD2.0: Towards a Reliable Multi-modal Sarcasm Detection System

2023-07-14 · Libo Qin, Shijue Huang, Qiguang Chen, Chenran Cai 외

Multi-modal sarcasm detection has attracted much recent attention. Nevertheless, the existing benchmark (MMSD) has some shortcomings that hinder the development of reliable multi-modal sarcasm detection system: (1) There…

Sarcasm Detection

Efficient Semantic Image Communication for Traffic Monitoring at the Edge

2026-04-14 · Damir Assylbek, Nurmukhammed Aitymbetov, Marko Ristin, Dimitrios Zorbas arxiv

Many visual monitoring systems operate under strict communication constraints, where transmitting full-resolution images is impractical and often unnecessary. In such settings, visual data is often used for object presen…

InterCLIP-MEP: Interactive CLIP and Memory-Enhanced Predictor for Multi-modal Sarcasm Detection

2024-06-24 · Junjie Chen, Hang Yu, Subin Huang, Sanmin Liu 외

Sarcasm in social media, often expressed through text-image combinations, poses challenges for sentiment analysis and intention mining. Current multi-modal sarcasm detection methods have been demonstrated to overly rely …

Sarcasm DetectionSentiment Analysis

MMSD-Net: Towards Multi-modal Stuttering Detection

2024-07-16 · Liangyu Nie, Sudarsana Reddy Kadiri, Ruchit Agrawal

Stuttering is a common speech impediment that is caused by irregular disruptions in speech production, affecting over 70 million people across the world. Standard automatic speech processing tools do not take speech ailm…

HCIG: A Hierarchical Cross-Modal Incongruity Graph Network for Multimodal Sarcasm and Cyberbullying Detection

2026-07-17 · Bhavana Verma, Priyanka Meel, Dinesh Kumar Vishwakarma arxiv

Multimodal sarcasm and cyberbullying detection remain challenging because the intended meaning often emerges from incongruity between textual and visual information rather than from either modality alone. Existing multim…

Multimodal Reasoning