paper-with-me

홈 › Papers

Lifting the Veil on Visual Information Flow in MLLMs: Unlocking Pathways to Faster Inference

2025-03-17 · CVPR 2025 1 · Hao Yin, Guangzong Si, Zilei Wang

Multimodal large language models (MLLMs) improve performance on vision-language tasks by integrating visual features from pre-trained vision encoders into large language models (LLMs). However, how MLLMs process and utilize visual information remains unclear. In this paper, a shift in the dominant flow of visual information is uncovered: (1) in shallow layers, strong interactions are observed between image tokens and instruction tokens, where most visual information is injected into instruction tokens to form cross-modal semantic representations; (2) in deeper layers, image tokens primarily interact with each other, aggregating the remaining visual information to optimize semantic representations within visual modality. Based on these insights, we propose Hierarchical Modality-Aware Pruning (HiMAP), a plug-and-play inference acceleration method that dynamically prunes image tokens at specific layers, reducing computational costs by approximately 65% without sacrificing performance. Our findings offer a new understanding of visual information processing in MLLMs and provide a state-of-the-art solution for efficient inference.

📄 PDF Abstract BibTeX arXiv:2503.13108

Code (1)

ustc-hyin/HiMAP 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

Pruning 설명 없음

Similar Papers 제목 키워드 기반

Suspicious Behavior Detection on Shoplifting Cases for Crime Prevention by Using 3D Convolutional Neural Networks

2020-04-30 · Guillermo A. Martínez-Mascorro, José R. Abreu-Pederzini, José C. Ortiz-Bayliss, Hugo Terashima-Marín

Crime generates significant losses, both human and economic. Every year, billions of dollars are lost due to attacks, crimes, and scams. Surveillance video camera networks are generating vast amounts of data, and the sur…

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems

2025-03-19 · Felix Chen, Hangjie Yuan, Yunqiu Xu, Tao Feng 외

Despite impressive performance across diverse tasks, Multimodal Large Language Models (MLLMs) have yet to fully demonstrate their potential in visual mathematical problem-solving, particularly in accurately perceiving an…

Mathematical Problem-Solving

Cross-modal Information Flow in Multimodal Large Language Models

2024-11-27 · CVPR 2025 1 · Zhi Zhang, Srishti Yadav, Fengze Han, Ekaterina Shutova

The recent advancements in auto-regressive multimodal large language models (MLLMs) have demonstrated promising progress for vision-language tasks. While there exists a variety of studies investigating the processing of …

Question AnsweringVisual Question Answering

LLaVAFlow: Preserving Latent Alignment Flow for Parameter-Efficient Multimodal Fine-Tuning

2026-08-27 · Muyao Yuan, Muyan Jiao, Jiangyong Ying, Weizhan Zhang 외 arxiv

While Multimodal Large Language Models (MLLMs) exhibit strong generalization, visual instruction tuning for downstream tasks inevitably causes catastrophic forgetting, impairing overall generalization. While existing met…

Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly

2024-06-15 · CVPR 2025 1 · Yexin Liu, Zhengyang Liang, Yueze Wang, Xianfeng Wu 외

Multimodal Large Language Models (MLLMs) have displayed remarkable performance in multi-modal tasks, particularly in visual comprehension. However, we reveal that MLLMs often generate incorrect answers even when they und…