paper-with-me

홈 › Papers

Modality Bias in LVLMs: Analyzing and Mitigating Object Hallucination via Attention Lens

2025-08-04 · Haohan Zheng, Zhenguo Zhang arxiv

Large vision-language models (LVLMs) have demonstrated remarkable multimodal comprehension and reasoning capabilities, but they still suffer from severe object hallucination. Previous studies primarily attribute the flaw to linguistic prior caused by the scale mismatch between visual encoders and large language models (LLMs) in LVLMs. Specifically, as current LVLMs are built upon LLMs, they tend to over-rely on textual prompts and internal knowledge of LLMs, generating descriptions inconsistent with visual cues. However, through an in-depth investigation of the hallucinated mechanisms, we empirically reveal a previously overlooked phenomenon: LVLMs may ignore not only visual information but also textual modality during hallucination, a behavior termed as modality bias, which indicates that LVLMs struggle to simultaneously attend to both visual and textual modalities, leading to fragmented understanding of user-provided instructions. Based on this observation, we propose a simple yet effective training-free method to mitigate object hallucination. Concretely, we intervene and adjust the attention weights of textual and visual tokens, balancing cross-modal compatibility for better alignment with user intentions. Furthermore, we adopt a contrastive decoding strategy to reduce the LVLM's overreliance on its parametric knowledge, synergistically enhancing our attention manipulation. Extensive experiments confirm the widespread presence of modality bias in LVLMs. Notably, our method effectively mitigates hallucination across multiple open-source LVLMs and benchmarks, highlighting its generalizability and efficacy.

📄 PDF Abstract BibTeX arXiv:2508.02419

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Analyzing and Mitigating Sycophancy in Large Vision-Language Models

2024-08-21 · Yunpu Zhao, Rui Zhang, Junbin Xiao, Changxin Ke 외

Large Vision-Language Models (LVLMs) have shown significant capability in vision-language understanding. However, one critical issue that persists in these models is sycophancy, which means models are unduly influenced b…

HallucinationPrompt Engineering

Analyzing and Mitigating Object Hallucination: A Training Bias Perspective

2025-08-06 · Yifan Li, Kun Zhou, Wayne Xin Zhao, Lei Fang 외 arxiv

As scaling up training data has significantly improved the general multimodal capabilities of Large Vision-Language Models (LVLMs), they still suffer from the hallucination issue, generating text that is inconsistent wit…

Analyzing and Mitigating Object Hallucination in Large Vision-Language Models

2023-10-01 · Yiyang Zhou, Chenhang Cui, Jaehong Yoon, Linjun Zhang 외

Large vision-language models (LVLMs) have shown remarkable abilities in understanding visual information with human languages. However, LVLMs still suffer from object hallucination, which is the problem of generating des…

HallucinationHallucination EvaluationObjectObject Hallucination

AFTER: Mitigating the Object Hallucination of LVLM via Adaptive Factual-Guided Activation Editing

2026-01-05 · Tianbo Wang, Yuqing Ma, Kewei Liao, Zhange Zhang 외 arxiv

Large Vision-Language Models (LVLMs) have achieved substantial progress in cross-modal tasks. However, due to language bias, LVLMs are susceptible to object hallucination, which can be primarily divided into category, at…

Mitigating Hallucinations in Large Vision-Language Models (LVLMs) via Language-Contrastive Decoding (LCD)

2024-08-06 · Avshalom Manevich, Reut Tsarfaty

Large Vision-Language Models (LVLMs) are an extension of Large Language Models (LLMs) that facilitate processing both image and text inputs, expanding AI capabilities. However, LVLMs struggle with object hallucinations d…

Object