Conditional Information Bottleneck for Multimodal Fusion: Overcoming Shortcut Learning in Sarcasm Detection
Multimodal sarcasm detection is a complex task that requires distinguishing subtle complementary signals across modalities while filtering out irrelevant information. Many advanced methods rely on learning shortcuts from datasets rather than extracting intended sarcasm-related features. However, our experiments show that shortcut learning impairs the model's generalization in real-world scenarios. Furthermore, we reveal the weaknesses of current modality fusion strategies for multimodal sarcasm detection through systematic experiments, highlighting the necessity of focusing on effective modality fusion for complex emotion recognition. To address these challenges, we construct MUStARD++$^{R}$ by removing shortcut signals from MUStARD++. Then, a Multimodal Conditional Information Bottleneck (MCIB) model is introduced to enable efficient multimodal fusion for sarcasm detection. Experimental results show that the MCIB achieves the best performance without relying on shortcut learning.
Code (0)
등록된 구현이 없습니다.
Tasks
Emotion RecognitionSarcasm DetectionSimilar Papers 제목 키워드 기반
Understanding and Overcoming Cross-modal Fusion Bias in Multimodal Anomaly Detection From A Fisher Information Perspective
Current advancements in Multimodal Anomaly Detection (MAD) are largely driven by enhancing multimodal fusion, particularly through the integration of RGB and Depth data for richer anomaly representation. However, less at…
Anomaly DetectionSkill-CMIB: Multimodal Agent Skill for Consistent Action via Conditional Multimodal Information Bottleneck
While LLM-based agents excel at planning and executing long action sequences, their execution often remains inconsistent across trials, limiting reliability. Consolidating agent consistency requires distilling trial-erro…
A Hidden Semantic Bottleneck in Conditional Embeddings of Diffusion Transformers
Diffusion Transformers have achieved state-of-the-art performance in class-conditional and multimodal generation, yet the structure of their learned conditional embeddings remains poorly understood. In this work, we pres…
multimodal generationAudio GenerationImage GenerationMultimodal Information Bottleneck: Learning Minimal Sufficient Unimodal and Multimodal Representations
Learning effective joint embedding for cross-modal data has always been a focus in the field of multimodal machine learning. We argue that during multimodal fusion, the generated multimodal embedding may be redundant, an…
Emotion RecognitionMultimodal Emotion RecognitionMultimodal Sentiment AnalysisSentiment AnalysisDenoising Bottleneck with Mutual Information Maximization for Video Multimodal Fusion
Video multimodal fusion aims to integrate multimodal signals in videos, such as visual, audio and text, to make a complementary prediction with multiple modalities contents. However, unlike other image-text multimodal ta…
DenoisingMultimodal Sentiment AnalysisSentiment Analysis