paper-with-me

홈 › Papers

Test-Time Adaptation of Vision-Language Models for Open-Vocabulary Semantic Segmentation

2025-05-28 · Mehrdad Noori, David Osowiechi, Gustavo Adolfo Vargas Hakim, Ali Bahri, Moslem Yazdanpanah, Sahar Dastani, Farzad Beizaee, Ismail Ben Ayed, Christian Desrosiers

Recently, test-time adaptation has attracted wide interest in the context of vision-language models for image classification. However, to the best of our knowledge, the problem is completely overlooked in dense prediction tasks such as Open-Vocabulary Semantic Segmentation (OVSS). In response, we propose a novel TTA method tailored to adapting VLMs for segmentation during test time. Unlike TTA methods for image classification, our Multi-Level and Multi-Prompt (MLMP) entropy minimization integrates features from intermediate vision-encoder layers and is performed with different text-prompt templates at both the global CLS token and local pixel-wise levels. Our approach could be used as plug-and-play for any segmentation network, does not require additional training data or labels, and remains effective even with a single test sample. Furthermore, we introduce a comprehensive OVSS TTA benchmark suite, which integrates a rigorous evaluation protocol, seven segmentation datasets, and 15 common corruptions, with a total of 82 distinct test scenarios, establishing a standardized and comprehensive testbed for future TTA research in open-vocabulary segmentation. Our experiments on this suite demonstrate that our segmentation-tailored method consistently delivers significant gains over direct adoption of TTA classification baselines.

📄 PDF Abstract BibTeX arXiv:2505.21844

Code (1)

dosowiechi/mlmp 공식 구현 pytorch

Tasks

image-classificationImage ClassificationOpen Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationSegmentationSemantic SegmentationTest-time Adaptation

Similar Papers 제목 키워드 기반

Effectiveness of Vision Language Models for Open-world Single Image Test Time Adaptation

2024-06-01 · Manogna Sreenivas, Soma Biswas

We propose a novel framework to address the real-world challenging task of Single Image Test Time Adaptation in an open and dynamic environment. We leverage large scale Vision Language Models like CLIP to enable real tim…

Contrastive LearningDomain AdaptationOut of Distribution (OOD) DetectionTest-time Adaptation

LQA: A Lightweight Quantized-Adaptive Framework for Vision-Language Models on the Edge

2026-02-08 · Xin Wang, Hong Jia, Hualin Zhou, Sheng Guang Wang 외 arxiv

Deploying Vision-Language Models (VLMs) on edge devices is challenged by resource constraints and performance degradation under distribution shifts. While test-time adaptation (TTA) can counteract such shifts, existing m…

Test-time Adaptation

ProtoDCS: Towards Robust and Efficient Open-Set Test-Time Adaptation for Vision-Language Models

2026-02-27 · Wei Luo, Yangfan Ou, Jin Deng, Zeshuai Deng 외 arxiv

Large-scale Vision-Language Models (VLMs) exhibit strong zero-shot recognition, yet their real-world deployment is challenged by distribution shifts. While Test-Time Adaptation (TTA) can mitigate this, existing VLM-based…

Test-time Adaptation

A Lost Opportunity for Vision-Language Models: A Comparative Study of Online Test-Time Adaptation for Vision-Language Models

2024-05-23 · Mario Döbler, Robert A. Marsden, Tobias Raichle, Bin Yang

In deep learning, maintaining model robustness against distribution shifts is critical. This work explores a broad range of possibilities to adapt vision-language foundation models at test-time, with a particular emphasi…

Image ClassificationPrompt EngineeringPrompt LearningTest-time Adaptation

CLIPTTA: Robust Contrastive Vision-Language Test-Time Adaptation

2025-07-18 · Marc Lafon, Gustavo Adolfo Vargas Hakim, Clément Rambour, Christian Desrosier 외 arxiv

Vision-language models (VLMs) like CLIP exhibit strong zero-shot capabilities but often fail to generalize under distribution shifts. Test-time adaptation (TTA) allows models to update at inference time without labeled d…

Test-time Adaptation