paper-with-me

홈 › Papers

Robust Multimodal Large Language Models Against Modality Conflict

2025-07-09 · Zongmeng Zhang, Wengang Zhou, Jie Zhao, Houqiang Li arxiv

Despite the impressive capabilities of multimodal large language models (MLLMs) in vision-language tasks, they are prone to hallucinations in real-world scenarios. This paper investigates the hallucination phenomenon in MLLMs from the perspective of modality conflict. Unlike existing works focusing on the conflicts between model responses and inputs, we study the inherent conflicts in inputs from different modalities that place MLLMs in a dilemma and directly lead to hallucinations. We formally define the modality conflict and construct a dataset named Multimodal Modality Conflict (MMMC) to simulate this phenomenon in vision-language tasks. Three methods based on prompt engineering, supervised fine-tuning, and reinforcement learning are proposed to alleviate the hallucination caused by modality conflict. Extensive experiments are conducted on the MMMC dataset to analyze the merits and demerits of these methods. Our results show that the reinforcement learning method achieves the best performance in mitigating the hallucination under modality conflict, while the supervised fine-tuning method shows promising and stable performance. Our work sheds light on the unnoticed modality conflict that leads to hallucinations and provides more insights into the robustness of MLLMs.

📄 PDF Abstract BibTeX arXiv:2507.07151

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningPrompt Engineering

Similar Papers 제목 키워드 기반

EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness

2026-05-01 · Yueru Sun, Yimeng Zhang, Haoyu Gu, Nuo Chen 외 arxiv

Multimodal Emotion Recognition (MER) is critical for interpreting real-world interactions. While Multimodal Large Language Models (MLLM) have shown promise in MER, their internal decision-making mechanisms under modality…

Multimodal Emotion Recognition

Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion Reasoning

2025-08-02 · Zhiyuan Han, Beier Zhu, Yanlong Xu, Peipei Song 외 arxiv

Despite their strong performance in multimodal emotion reasoning, existing Multimodal Large Language Models (MLLMs) often overlook the scenarios involving emotion conflicts, where emotional cues from different modalities…

Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis

2023-10-09 · Haoyu Zhang, Yu Wang, Guanghao Yin, Kejun Liu 외

Though Multimodal Sentiment Analysis (MSA) proves effective by utilizing rich information from multiple sources (e.g., language, video, and audio), the potential sentiment-irrelevant and conflicting information across mo…

Multimodal Sentiment AnalysisSentiment Analysis

Seeing Sound, Hearing Sight: Uncovering Modality Bias and Conflict of AI models in Sound Localization

2025-05-16 · Yanhao Jia, Ji Xie, S Jivaganesh, Hao Li 외

Imagine hearing a dog bark and turning toward the sound only to see a parked car, while the real, silent dog sits elsewhere. Such sensory conflicts test perception, yet humans reliably resolve them by prioritizing sound …

Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

2024-10-04 · Tinghui Zhu, Qin Liu, Fei Wang, Zhengzhong Tu 외

Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities for capturing and reasoning over multimodal inputs. However, these models are prone to parametric knowledge conflicts, which arise from incon…