paper-with-me

Papers

VIP: Visual-guided Prompt Evolution for Efficient Dense Vision-Language Inference

2026-05-12 · Hao Zhu, Shuo Jin, Wenbin Liao, Jiayu Xiao, Yan Zhu, Siyue Yu, Feng Dai arxiv

Pursuing training-free open-vocabulary semantic segmentation in an efficient and generalizable manner remains challenging due to the deep-seated spatial bias in CLIP. To overcome the limitations of existing solutions, this work moves beyond the CLIP-based paradigm and harnesses the recent spatially-aware dino$.$txt framework to facilitate more efficient and high-quality dense prediction. While dino$.$txt exhibits robust spatial awareness, we find that the semantic ambiguity of text queries gives rise to severe mismatch within its dense cross-modal interactions. To address this, we introduce Visual-guided Prompt evolution (VIP) to rectify the semantic expressiveness of text queries in dino$.$txt, unleashing its potential for fine-grained object perception. Towards this end, VIP integrates alias expansion with a visual-guided distillation mechanism to mine valuable semantic cues, which are robustly aggregated in a saliency-aware manner to yield a high-fidelity prediction. Extensive evaluations demonstrate that VIP: 1. surpasses the top-leading methods by 1.4%-8.4% average mIoU, 2. generalizes well to diverse challenging domains, and 3. requires marginal inference time and memory overhead.

📄 PDF Abstract BibTeX arXiv:2605.12325

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic Segmentation

Similar Papers 제목 키워드 기반

VisFocus: Prompt-Guided Vision Encoders for OCR-Free Dense Document Understanding

2024-07-17 · Ofir Abramovich, Niv Nayman, Sharon Fogel, Inbal Lavi 외

In recent years, notable advancements have been made in the domain of visual document understanding, with the prevailing architecture comprising a cascade of vision and language models. The text component can either be e…

document understandingOptical Character Recognition (OCR)

Aligning Medical Images with General Knowledge from Large Language Models

2024-08-31 · Xiao Fang, Yi Lin, Dong Zhang, Kwang-Ting Cheng 외

Pre-trained large vision-language models (VLMs) like CLIP have revolutionized visual representation learning using natural language as supervisions, and demonstrated promising generalization ability. In this work, we pro…

General KnowledgeMedical Image AnalysisPrompt LearningRepresentation Learning+1

Guiding Evolution of Artificial Life Using Vision-Language Models

2025-09-26 · Nikhil Baid, Hannah Erlebach, Paul Hellegouarch, Frederico Wieser arxiv

Foundation models (FMs) have recently opened up new frontiers in the field of artificial life (ALife) by providing powerful tools to automate search through ALife simulations. Previous work aligns ALife simulations with …

PromptMoE: Generalizable Zero-Shot Anomaly Detection via Visually-Guided Prompt Mixtures

2025-11-22 · Yuheng Shao, Lizhang Wang, Changhao Li, Peixian Chen 외 arxiv

Zero-Shot Anomaly Detection (ZSAD) aims to identify and localize anomalous regions in images of unseen object classes. While recent methods based on vision-language models like CLIP show promise, their performance is con…

Prompt EngineeringAnomaly Detection

DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting

2021-12-02 · CVPR 2022 1 · Yongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang 외

Recent progress has shown that large-scale pre-training using contrastive image-text pairs can be a promising alternative for high-quality visual representation learning from natural language supervision. Benefiting from…

Image-text matchingInstance SegmentationLanguage Modellingobject-detection+5