paper-with-me

홈 › Papers

Supporting Vision-Language Model Inference with Causality-pruning Knowledge Prompt

2022-05-23 · Jiangmeng Li, Wenyi Mo, Wenwen Qiang, Bing Su, Changwen Zheng

Vision-language models are pre-trained by aligning image-text pairs in a common space so that the models can deal with open-set visual concepts by learning semantic information from textual labels. To boost the transferability of these models on downstream tasks in a zero-shot manner, recent works explore generating fixed or learnable prompts, i.e., classification weights are synthesized from natural language describing task-relevant categories, to reduce the gap between tasks in the training and test phases. However, how and what prompts can improve inference performance remains unclear. In this paper, we explicitly provide exploration and clarify the importance of including semantic information in prompts, while existing prompt methods generate prompts without exploring the semantic information of textual labels. A challenging issue is that manually constructing prompts, with rich semantic information, requires domain expertise and is extremely time-consuming. To this end, we propose Causality-pruning Knowledge Prompt (CapKP) for adapting pre-trained vision-language models to downstream image recognition. CapKP retrieves an ontological knowledge graph by treating the textual label as a query to explore task-relevant semantic information. To further refine the derived semantic information, CapKP introduces causality-pruning by following the first principle of Granger causality. Empirically, we conduct extensive evaluations to demonstrate the effectiveness of CapKP, e.g., with 8 shots, CapKP outperforms the manual-prompt method by 12.51% and the learnable-prompt method by 1.39% on average, respectively. Experimental analyses prove the superiority of CapKP in domain generalization compared to benchmark approaches.

📄 PDF Abstract BibTeX arXiv:2205.11100

Code (0)

등록된 구현이 없습니다.

Tasks

Domain GeneralizationLanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Object-Centric Vision Token Pruning for Vision Language Models

2025-11-25 · Guangyuan Li, Rongzhen Zhao, Jinhong Deng, Yanbo Wang 외 arxiv

In Vision Language Models (VLMs), vision tokens are quantity-heavy yet information-dispersed compared with language tokens, thus consume too much unnecessary computation. Pruning redundant vision tokens for high VLM infe…

Deriving Coding-Specific Sub-Models from LLMs using Resource-Efficient Pruning

2025-01-09 · Laura Puccioni, Alireza Farshin, Mariano Scazzariello, Changjie Wang 외

Large Language Models (LLMs) have demonstrated their exceptional performance in various complex code generation tasks. However, their broader adoption is limited by significant computational demands and high resource req…

Code Generation

LVPruning: An Effective yet Simple Language-Guided Vision Token Pruning Approach for Multi-modal Large Language Models

2025-01-23 · Yizheng Sun, Yanze Xin, Hao Li, Jingyuan Sun 외

Multi-modal Large Language Models (MLLMs) have achieved remarkable success by integrating visual and textual modalities. However, they incur significant computational overhead due to the large number of vision tokens pro…

AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference with Dynamical Text Guidance

2025-08-08 · Weichen Zhang, Zhui Zhu, Ningbo Li, Shilong Tao 외 arxiv

Vision-language models (VLMs) have achieved impressive performance on multimodal reasoning tasks such as visual question answering, image captioning and so on, but their inference cost remains a significant challenge due…

Visual Question AnsweringMultimodal ReasoningImage Captioning

Do Large Language Models Show Biases in Causal Learning? Insights from Contingency Judgment

2025-10-15 · María Victoria Carro, Denise Alejandra Mester, Francisca Gauna Selasco, Giovanni Franco Gabriel Marraffini 외 arxiv

Causal learning is the cognitive process of developing the capability of making causal inferences based on available information, often guided by normative principles. This process is prone to errors and biases, such as …