paper-with-me

홈 › Papers

Investigating Anisotropy in Visual Grounding under Controlled Counterfactual Perturbations

2026-05-09 · Gabriele Lombardo, Luigi Maiorana, Liliana Lo Presti, Marco La Cascia arxiv

Visual Grounding benchmarks assume that the object described by a referring expression is always present in the image, and grounding models are therefore rarely evaluated under semantically mismatched captions. In such cases, models frequently exhibit approximation behavior, producing a plausible bounding box that satisfies only part of the expression (\eg, preserving the original object while ignoring modified contextual cues). Because mismatched captions represent realistic edge cases, this behavior compromises reliability and raises concerns from an explainability perspective. Identifying its underlying causes is thus essential for improving model faithfulness and interpretability. Adopting a mechanistic interpretability viewpoint, this work examines whether embedding anisotropy contributes to counterfactual failures. A similarity-controlled counterfactual caption generation protocol is introduced to systematically perturb object or contextual components within predefined embedding similarity intervals, enabling a fine-grained analysis of grounding behavior as a function of alignment. Experiments on two Transformer-based models with markedly different embedding geometries (BERT-based TransVG and CLIP-based SwimVG) reveal no meaningful correlation between cosine similarity and approximation. These findings suggest that anisotropy alone does not account for counterfactual errors, and that robustness requires investigating finer-grained geometric properties of the embedding space.

📄 PDF Abstract BibTeX arXiv:2605.09090

Code (0)

등록된 구현이 없습니다.

Tasks

Referring ExpressionVisual Grounding

Similar Papers 제목 키워드 기반

Significance of Natural Scene Statistics in Understanding the Anisotropies of Perceptual Filling-in at the Blind Spot

2017-01-12

Psychophysical experiments reveal our horizontal preference in perceptual filling-in at the blind spot. On the other hand, vertical preference is exhibited in the case of tolerance in filling-in. What causes this anisotr…

Investigating Compositional Challenges in Vision-Language Models for Visual Grounding

2024-01-01 · CVPR 2024 1 · Yunan Zeng, Yan Huang, Jinjin Zhang, Zequn Jie 외

Pre-trained vision-language models (VLMs) have achieved high performance on various downstream tasks which have been widely used for visual grounding tasks in a weakly supervised manner. However despite the performan…

AttributeRelationVisual Grounding

Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs

2026-07-04 · Anas Zafar, Leema Krishna Murali, Siddhant Bharadwaj, Ashish Vashist 외 arxiv

Large vision language models (VLMs) report strong accuracy on medical question-answering, yet it remains unclear whether they reason from visual evidence or exploit textual shortcuts. We introduce a counterfactual evalua…

A Novel Field-Free SOT Magnetic Tunnel Junction With Local VCMA-Induced Switching

2023-12-24 · Rui Zhou, Haiyang Zhang, Hao Wang, Jin He 외

By integrating the local voltage-controlled magnetic anisotropy (VCMA) effect, Dzyaloshinskii-Moriya interaction (DMI) effect, and spin-orbit torque (SOT) effect, we propose a novel device structure for field-free magnet…

Attending Self-Attention: A Case Study of Visually Grounded Supervision in Vision-and-Language Transformers

2021-08-01 · ACL 2021 5 · Jules Samaran, Noa Garcia, Mayu Otani, Chenhui Chu 외

The impressive performances of pre-trained visually grounded language models have motivated a growing body of research investigating what has been learned during the pre-training. As a lot of these models are based on Tr…

Language ModelingLanguage ModellingVisual Grounding