paper-with-me

홈 › Papers

Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens

2024-11-23 · CVPR 2025 1 · Zhangqi Jiang, Junkai Chen, Beier Zhu, Tingjin Luo, Yankun Shen, Xu Yang

Hallucinations in Large Vision-Language Models (LVLMs) significantly undermine their reliability, motivating researchers to explore the causes of hallucination. However, most studies primarily focus on the language aspect rather than the visual. In this paper, we address how LVLMs process visual information and whether this process causes hallucination. Firstly, we use the attention lens to identify the stages at which LVLMs handle visual data, discovering that the middle layers are crucial. Moreover, we find that these layers can be further divided into two stages: ''visual information enrichment'' and ''semantic refinement'' which respectively propagate visual data to object tokens and interpret it through text. By analyzing attention patterns during the visual information enrichment stage, we find that real tokens consistently receive higher attention weights than hallucinated ones, serving as a strong indicator of hallucination. Further examination of multi-head attention maps reveals that hallucination tokens often result from heads interacting with inconsistent objects. Based on these insights, we propose a simple inference-time method that adjusts visual attention by integrating information across various heads. Extensive experiments demonstrate that this approach effectively mitigates hallucinations in mainstream LVLMs without additional training costs. Code is available at https://github.com/ZhangqiJiang07/middle_layers_indicating_hallucinations.

📄 PDF Abstract BibTeX arXiv:2411.16724

Code (1)

zhangqijiang07/middle_layers_indicating_hallucinations 공식 구현 pytorch

Tasks

Hallucination

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Focus 설명 없음
Multi-Head Attention 설명 없음

Similar Papers 제목 키워드 기반

Unveiling the Reasoning Process of Large Language Models

2026-03-31 · Junjie Zhang, Zhen Shen, Xisong Dong, Gang Xiong arxiv

Large language models often reason beyond surface tokens, but the internal stage at which token-level information becomes abstract relational structure remains unclear. We investigate this question by analyzing how atten…

Enhancing Non-English Capabilities of English-Centric Large Language Models through Deep Supervision Fine-Tuning

2025-03-03 · Wenshuai Huo, Xiaocheng Feng, Yichong Huang, Chengpeng Fu 외

Large language models (LLMs) have demonstrated significant progress in multilingual language understanding and generation. However, due to the imbalance in training data, their capabilities in non-English languages are l…

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Pruning

2026-08-04 · Yuyao Sun, Tao Deng, Shuang Li, Deqing Wang 외 arxiv

Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language tasks, but their efficiency is limited by the cost of processing numerous visual tokens. Visual token pruning can reduce …

Unfair Alignment: Examining Safety Alignment Across Vision Encoder Layers in Vision-Language Models

2024-11-06 · Saketh Bachu, Erfan Shayegani, Trishna Chakraborty, Rohit Lal 외

Vision-language models (VLMs) have improved significantly in multi-modal tasks, but their more complex architecture makes their safety alignment more challenging than the alignment of large language models (LLMs). In thi…

Safety Alignment

Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs

2025-10-15 · Minji Kim, Taekyung Kim, Bohyung Han arxiv

Video Large Language Models (VideoLLMs) extend the capabilities of vision-language models to spatiotemporal inputs, enabling tasks such as video question answering (VideoQA). Despite recent advances in VideoLLMs, their i…

Video Question Answering