paper-with-me

홈 › Papers

Gram-Anchored Prompt Learning for Vision-Language Models via Second-Order Statistics

2026-04-05 · Minglei Chen, Weilong Wang, Jiang Duan, Ye Deng arxiv

Parameter-efficient prompt learning has become the de facto standard for adapting Vision-Language Models (VLMs) to downstream tasks. Existing approaches predominantly focus on aligning text prompts with first-order visual features (i.e., spatial feature maps). While effective for fine-grained semantic discrimination, we argue that relying solely on first-order information is insufficient for robust adaptation, as these spatially entangled features are highly susceptible to domain shifts and local noise. In this work, we propose \textbf{Gram-Anchored Prompt Learning (GAPL)} for Vision-Language Models via Second-Order Statistics, a framework that synergizes local semantic alignment with global structural consistency. Methodologically, we introduce an additional second-order statistical stream via \textbf{Gram matrices} that augments the standard first-order spatial interaction. By anchoring prompts to these second-order priors, our approach enables language representations to dynamically adapt to statistical distribution shifts across diverse domains. Extensive experiments indicate the effectiveness of the second-order features, and show compelling performances of GAPL on various benchmarks.

📄 PDF Abstract BibTeX arXiv:2604.03980

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Prompt-Anchored Vision-Text Distillation for Lifelong Person Re-identification

2026-05-06 · Wen Wen, Hao Chen, Shiliang Zhang arxiv

Lifelong person re-identification (LReID) aims to train a generalizable model with sequentially collected data. However, such models often suffer from semantic drift, limited adaptability, and catastrophic forgetting as …

Person Re-IdentificationDomain Generalization

An Empirical Audit of k-NAF Budget Accounting for Anchored Decoding

2026-05-27 · J. Vijayavallabh arxiv

We empirically audit the k-NAF budget-accounting mechanism in Anchored Decoding using (i) a fixed, class-stratified workload (approximately 8,500 randomized executions across six prompt classes) and (ii) an adaptive prom…

K-MaT: Knowledge-Anchored Manifold Transport for Cross-Modal Prompt Learning in Medical Imaging

2026-03-06 · Jiajun Zeng, Shadi Albarqouni arxiv

Large-scale biomedical vision-language models (VLMs) adapted on high-end imaging (e.g., CT) often fail to transfer to frontline low-end modalities (e.g., radiography), collapsing into modality-specific shortcuts. We prop…

Prompting Large Pre-trained Vision-Language Models For Compositional Concept Learning

2022-11-09 · Guangyue Xu, Parisa Kordjamshidi, Joyce Chai

This work explores the zero-shot compositional learning ability of large pre-trained vision-language models(VLMs) within the prompt-based learning framework and propose a model (\textit{PromptCompVL}) to solve the compos…

Zero-Shot Learning

Sequential Planning via Anchored Robotic Keypoints

2026-06-29 · Bryce Grant, Aryeh Rothenberg, Logan Senning, Zonghe Chua 외 arxiv

We present Sequential Planning via Anchored Robotic Keypoints, SPARK, a training-free neurosymbolic manipulation system that reaches 43.7% on six LIBERO-PRO position \& task cells, more than doubling CaP-Agent0 and Visio…