paper-with-me

홈 › Papers

Language-Guided Invariance Probing of Vision-Language Models

2025-11-17 · Jae Joong Lee arxiv

Recent vision-language models (VLMs) such as CLIP, OpenCLIP, EVA02-CLIP and SigLIP achieve strong zero-shot performance, but it is unclear how reliably they respond to controlled linguistic perturbations. We introduce Language-Guided Invariance Probing (LGIP), a benchmark that measures (i) invariance to meaning-preserving paraphrases and (ii) sensitivity to meaning-changing semantic flips in image-text matching. Using 40k MS COCO images with five human captions each, we automatically generate paraphrases and rule-based flips that alter object category, color or count, and summarize model behavior with an invariance error, a semantic sensitivity gap and a positive-rate statistic. Across nine VLMs, EVA02-CLIP and large OpenCLIP variants lie on a favorable invariance-sensitivity frontier, combining low paraphrase-induced variance with consistently higher scores for original captions than for their flipped counterparts. In contrast, SigLIP and SigLIP2 show much larger invariance error and often prefer flipped captions to the human descriptions, especially for object and color edits. These failures are largely invisible to standard retrieval metrics, indicating that LGIP provides a model-agnostic diagnostic for the linguistic robustness of VLMs beyond conventional accuracy scores.

📄 PDF Abstract BibTeX arXiv:2511.13494

Code (0)

등록된 구현이 없습니다.

Tasks

Image-text matching

Similar Papers 제목 키워드 기반

Blind to Position, Biased in Language: Probing Mid-Layer Representational Bias in Vision-Language Encoders for Zero-Shot Language-Grounded Spatial Understanding

2025-09-27 · Na Min An, Inha Kang, Minhyun Lee, Hyunjung Shim arxiv

Vision-Language Encoders (VLEs) are widely adopted as the backbone of zero-shot referring image segmentation (RIS), enabling text-guided localization without task-specific training. However, prior works underexplored the…

Image SegmentationImage Retrieval

Probing Semantic Alignment, Lexical Invariance, and Syntactic Influence in LLM Metaphor Processing

2025-10-05 · Fengying Ye, Shanshan Wang, Lidia S. Chao, Derek F. Wong arxiv

Large language models (LLMs) achieve strong performance on metaphor detection and interpretation tasks, yet it remains unclear what such behavioral success reveals about metaphor processing. We present a diagnostic analy…

Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability

2026-02-10 · Aaditya Vikram Prasad, Connor Watts, Jack Merullo, Dhruvil Gala 외 arxiv

Language models trained on large-scale datasets have been shown to learn features that encode abstract concepts such as factuality or intent. Such features are traditionally used for test-time monitoring or steering. We …

Reinforcement Learning

Probing The Linguistic Capacity of Pre-Trained Vision-Language Models

2022-01-16 · ACL ARR January 2022 1 · Anonymous

How do recent vision-language pre-trained models compare against language-specific pre-trained models on common linguistic tasks? In this paper, we assess this in a probing setting. Our results suggest that different mul…

Transitive Vision-Language Prompt Learning for Domain Generalization

2024-04-29 · Liyuan Wang, Yan Jin, Zhen Chen, Jinlin Wu 외

The vision-language pre-training has enabled deep models to make a huge step forward in generalizing across unseen domains. The recent learning method based on the vision-language pre-training model is a great tool for d…

Domain GeneralizationPrompt Learning