paper-with-me

홈 › Papers

Reallocating Attention Across Layers to Reduce Multimodal Hallucination

2025-10-11 · Haolang Lu, Bolun Chu, WeiYe Fu, Guoshun Nan, Junning Liu, Minghui Pan, Qiankun Li, Yi Yu, Hua Wang, Kun Wang arxiv

Multimodal large reasoning models (MLRMs) often suffer from hallucinations that stem not only from insufficient visual grounding but also from imbalanced allocation between perception and reasoning processes. Building upon recent interpretability findings suggesting a staged division of attention across layers, we analyze how this functional misalignment leads to two complementary failure modes: perceptual bias in shallow layers and reasoning drift in deeper layers. To alleviate these issues, we propose Functional Head Identification and Class-Conditioned Rescaling , a lightweight, training-free plugin that identifies perception- and reasoning-oriented heads and adaptively rebalances their layerwise contributions. Our method improves reasoning consistency and visual faithfulness without retraining or any architectural modification. Evaluations across three representative MLRMs and five multimodal reasoning benchmarks show an average 4.2% point gain, with less than 1% additional computation and only 9% baseline latency. Beyond empirical improvements, our study provides an interpretable perspective on regulating cross-layer functional dynamics to enhance the reliability of multimodal reasoning.

📄 PDF Abstract BibTeX arXiv:2510.10285

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal ReasoningVisual Grounding

Similar Papers 제목 키워드 기반

GuardAlign: Test-time Safety Alignment in Multimodal Large Language Models

2026-02-27 · Xingyu Zhu, Beier Zhu, Junfeng Fang, Shuo Wang 외 arxiv

Large vision-language models (LVLMs) have achieved remarkable progress in vision-language reasoning tasks, yet ensuring their safety remains a critical challenge. Recent input-side defenses detect unsafe images with CLIP…

Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping

2025-10-10 · Dwip Dalal, Gautam Vashishtha, Utkarsh Mishra, Jeonghwan Kim 외 arxiv

Multimodal large language models (MLLMs) often miss small details and spatial relations in cluttered scenes, leading to errors in fine-grained perceptual grounding. We introduce AttWarp, a lightweight method that allocat…

Efficient Token Pruning for LLaDA-V

2026-01-28 · Zhewen Wan, Tianchen Song, Chen Lin, Zhiyong Zhao 외 arxiv

Diffusion-based large multimodal models, such as LLaDA-V, have demonstrated impressive capabilities in vision-language understanding and generation. However, their bidirectional attention mechanism and diffusion-style it…

$\mathcal{V}isi\mathcal{P}runer$: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMs

2025-10-20 · Yingqi Fan, Anhao Zhao, Jinlan Fu, Junlong Tong 외 arxiv

Multimodal Large Language Models (MLLMs) have achieved strong performance across vision-language tasks, but suffer from significant computational overhead due to the quadratic growth of attention computations with the nu…

Vision Token Reduction via Attention-Driven Self-Compression for Efficient Multimodal Large Language Models

2026-02-13 · Omer Faruk Deniz, Ruiyu Mao, Ruochen Li, Yapeng Tian 외 arxiv

Multimodal Large Language Models (MLLMs) incur significant computational cost from processing numerous vision tokens through all LLM layers. Prior pruning methods operate either before the LLM, limiting generality due to…