paper-with-me

홈 › Papers

ID-Selection: Importance-Diversity Based Visual Token Selection for Efficient LVLM Inference

2026-04-07 · Zhaohong Huang, Wenjing Liu, Yuxin Zhang, Fei Chao, Rongrong Ji arxiv

Recent advances have explored visual token pruning to accelerate the inference of large vision-language models (LVLMs). However, existing methods often struggle to balance token importance and diversity: importance-based methods tend to retain redundant tokens, whereas diversity-based methods may overlook informative ones. This trade-off becomes especially problematic under high reduction ratios, where preserving only a small subset of visual tokens is critical. To address this issue, we propose ID-Selection, a simple yet effective token selection strategy for efficient LVLM inference. The key idea is to couple importance estimation with diversity-aware iterative selection: each token is first assigned an importance score, after which high-scoring tokens are selected one by one while the scores of similar tokens are progressively suppressed. In this way, ID-Selection preserves informative tokens while reducing redundancy in a unified selection process. Extensive experiments across 5 LVLM backbones and 16 main benchmarks demonstrate that ID-Selection consistently achieves superior performance and efficiency, especially under extreme pruning ratios. For example, on LLaVA-1.5-7B, ID-Selection prunes 97.2% of visual tokens, retaining only 16 tokens, while reducing inference FLOPs by over 97% and preserving 91.8% of the original performance, all without additional training.

📄 PDF Abstract BibTeX arXiv:2604.05601

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Stage-adaptive Token Selection for Efficient Omni-modal LLMs

2026-05-19 · Zijie Xin, Jie Yang, Ruixiang Zhao, Tianyi Wang 외 arxiv

Omni-modal large language models (om-LLMs) achieve unified audio-visual understanding by encoding video and audio into temporally aligned token sequences interleaved at the window level. However, processing these dense n…

Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs

2026-07-25 · Jouwon Song, Woohyeong Kim, Kyeongbo Kong arxiv

Recent high-resolution Multimodal Large Language Models (MLLMs) generate thousands of visual tokens per input, leading to a visual token explosion that introduces severe latency bottlenecks. While token pruning mitigates…

AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning

2026-07-08 · Kyuan Oh, Bumsoo Kim arxiv

Large vision-language models incur substantial inference costs because high-resolution inputs introduce thousands of visual tokens, many of which are redundant for a given query. Existing pruning methods often combine qu…

Beyond Surrogate Gradients: Fully Differentiable Token Pruning for Vision-Language Models

2026-05-27 · Landi He, Mingde Yao, Shawn Young, Lijian Xu arxiv

Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. Existing methods typically rely on Gumbel-Softmax to approximate discrete selection during trainin…

Continuous Control

Moving Beyond Diversity: Visual Token Pruning as Subspace Reconstruction for Efficient VLMs

2026-06-17 · Jaeyeon Lee, Shunjie Wen, Dong-Wan Choi arxiv

Despite their remarkable performance, Vision Language Models (VLMs) incur substantial computational overhead due to the large number of visual tokens. While diversity maximization has become a dominant strategy for token…