paper-with-me

홈 › Papers

Conical Visual Concentration for Efficient Large Vision-Language Models

2025-01-01 · CVPR 2025 1 · Long Xing, Qidong Huang, Xiaoyi Dong, Jiajie Lu, Pan Zhang, Yuhang Zang, Yuhang Cao, Conghui He, Jiaqi Wang, Feng Wu, Dahua Lin

In large vision-language models (LVLMs), images serve as inputs that carry a wealth of information. As the idiom "A picture is worth a thousand words" implies, representing a single image in current LVLMs can require hundreds or even thousands of tokens. This results in significant computational costs, which grow quadratically as input image resolution increases, thereby severely impacting the efficiency. Previous approaches have attempted to reduce the number of image tokens either before or within the early layers of LVLMs. However, these strategies inevitably result in the loss of crucial image information. To address this challenge, we conduct an empirical study revealing that all visual tokens are necessary for LVLMs in the shallow layers, and token redundancy progressively increases in the deeper layers.To this end, we propose ViCo, a conical-style visual concentration strategy for LVLMs to boost their efficiency in both training and inference with neglectable performance loss. Specifically, we partition the LVLM into several stages and drop part of the image tokens at the end of each stage with a pre-defined ratio. The dropping is based on a lightweight similarity calculation with a negligible time overhead. Extensive experiments demonstrate that ViCo can achieve over 40% training time reduction and 55% inference FLOPs acceleration on leading LVLMs like LLaVA-NeXT, maintaining comparable multi-modal performance. Besides, ViCo can also serve as a plug-and-play strategy to accelerate inference in a free way, with better performance and lower inference cost than counterparts.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Quantum-inspired Algorithm for General Minimum Conical Hull Problems

2019-07-16 · Yuxuan Du, Min-Hsiu Hsieh, Tongliang Liu, DaCheng Tao

A wide range of fundamental machine learning tasks that are addressed by the maximum a posteriori estimation can be reduced to a general minimum conical hull problem. The best-known solution to tackle general minimum con…

Near-separable Non-negative Matrix Factorization with $\ell_1$- and Bregman Loss Functions

2013-12-27 · Abhishek Kumar, Vikas Sindhwani

Recently, a family of tractable NMF algorithms have been proposed under the assumption that the data matrix satisfies a separability condition Donoho & Stodden (2003); Arora et al. (2012). Geometrically, this condition r…

HCNQA: Enhancing 3D VQA with Hierarchical Concentration Narrowing Supervision

2025-07-02 · Shengli Zhou, Jianuo Zhu, Qilin Huang, Fangjing Wang 외 arxiv

3D Visual Question-Answering (3D VQA) is pivotal for models to perceive the physical world and perform spatial reasoning. Answer-centric supervision is a commonly used training method for 3D VQA models. Many models that …

Spatial Reasoning

LASER: A Corrective Lens for LVLMs via Visual Attention Preservation and Sink Suppression

2026-07-02 · Bowen Yuan, Zijian Wang, Yadan Luo, Shijie Wang 외 arxiv

Large vision-language models (LVLMs) exhibit strong reasoning ability but suffer from visual forgetting during long-horizon decoding, where attention progressively drifts away from visual evidence. Existing methods large…

Visual Grounding

Divide-and-Conquer Learning by Anchoring a Conical Hull

2014-06-22 · NeurIPS 2014 12 · Tianyi Zhou, Jeff Bilmes, Carlos Guestrin

We reduce a broad class of machine learning problems, usually addressed by EM or sampling, to the problem of finding the $k$ extremal rays spanning the conical hull of a data point set. These $k$ "anchors" lead to a glob…

Clustering