paper-with-me

홈 › Papers

Make Your LVLM KV Cache More Lightweight

2026-05-01 · Xihao Chen, Yangyang Guo, Roger Zimmermann arxiv

Key-Value (KV) cache has become a de facto component of modern Large Vision-Language Models (LVLMs) for inference. While it enhances decoding efficiency in Large Language Models (LLMs), its direct adoption in LVLMs introduces substantial GPU memory overhead due to the large number of vision tokens processed during the prefill stage. To tackle this problem, we propose LightKV, a novel approach that reduces KV cache size by exploiting the redundancy among vision-token embeddings. Guided by text prompts, LightKV employs cross-modality message passing to aggregate informative messages across vision tokens and progressively compress them during prefill. This prompt-aware guidance distinguishes our method from prior vision-only compression strategies. We evaluate LightKV on eight open-source LVLMs across eight public benchmark datasets, e.g., MME and SeedBench. Experimental results demonstrate that with only 55% of the original vision tokens, LightKV (a) halves the vision-token KV cache size, (b) reduces computation by up to 40%, and (c) preserves general-purpose performance while significantly outperforming existing baselines.

📄 PDF Abstract BibTeX arXiv:2605.00789

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Edit-Your-Interest: Efficient Video Editing via Feature Most-Similar Propagation

2025-10-15 · Yi Zuo, Zitao Wang, Lingling Li, Xu Liu 외 arxiv

Text-to-image (T2I) diffusion models have recently demonstrated significant progress in video editing. However, existing video editing methods are severely limited by their high computational overhead and memory consumpt…

Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models

2025-10-23 · Xuyang Liu, Xiyan Gui, Yuchao Zhang, Linfeng Zhang arxiv

Recent large vision-language models (LVLMs) demonstrate remarkable capabilities in processing extended multi-modal sequences, yet the resulting key-value (KV) cache expansion creates a critical memory bottleneck that fun…

AirCache: Activating Inter-modal Relevancy KV Cache Compression for Efficient Large Vision-Language Model Inference

2025-03-31 · Kai Huang, Hao Zou, Bochen Wang, Ye Xi 외

Recent advancements in Large Visual Language Models (LVLMs) have gained significant attention due to their remarkable reasoning capabilities and proficiency in generalization. However, processing a large number of visual…

Language ModelingLanguage Modelling

Does Your Vision-Language Model Get Lost in the Long Video Sampling Dilemma?

2025-03-16 · Tianyuan Qu, Longxiang Tang, Bohao Peng, Senqiao Yang 외

The rise of Large Vision-Language Models (LVLMs) has significantly advanced video understanding. However, efficiently processing long videos remains a challenge due to the ``Sampling Dilemma'': low-density sampling risks…

Language ModelingLanguage ModellingVideo Understanding

Efficient Inference of Vision Instruction-Following Models with Elastic Cache

2024-07-25 · Zuyan Liu, Benlin Liu, Jiahui Wang, Yuhao Dong 외

In the field of instruction-following large vision-language models (LVLMs), the efficient deployment of these models faces challenges, notably due to the high memory demands of their key-value (KV) caches. Conventional c…

Instruction FollowingText Generation