paper-with-me

홈 › Papers

Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency

2026-05-13 · Ziqi Wen, Parsa Madinei, Miguel P. Eckstein arxiv

Evaluating whether large vision-language models (VLMs) align with human perception for high-level semantic scene comprehension remains a challenge. Traditional white-box interpretability methods are inapplicable to closed-source architectures and passive metrics fail to isolate causal features. We introduce Counterfactual Semantic Saliency (CSS). This black-box, model-agnostic framework quantifies the importance of objects by measuring the semantic shift induced by their causal ablation from a scene. To evaluate AI-human semantic alignment, we tested prominent VLMs against a human psychophysics baseline comprising 16,289 valid responses across 307 complex natural scenes and 1,306 high-fidelity counterfactual variants. Our analysis reveals a pervasive scene comprehension gap: models exhibit an overreliance (relative to humans) on large objects (size bias), objects at the center of the image (center bias), and high saliency objects. In contrast, models rely less on people in the scenes than our human participants to describe the images. A model's size bias is a primary driver explaining variations in model-human semantic divergence. Code and data will be available at https://github.com/starsky77/Counterfactual-Semantic-Saliency.

📄 PDF Abstract BibTeX arXiv:2605.13047

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning

2024-05-02 · Shihao Wang, Zhiding Yu, Xiaohui Jiang, Shiyi Lan 외

The advances in vision-language models (VLMs) have led to a growing interest in autonomous driving to leverage their strong reasoning capabilities. However, extending these capabilities from 2D to full 3D understanding i…

Autonomous DrivingcounterfactualCounterfactual ReasoningDecision Making+3

OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning

2025-04-06 · CVPR 2025 1 · Shihao Wang, Zhiding Yu, Xiaohui Jiang, Shiyi Lan 외

The advances in vision-language models (VLMs) have led to a growing interest in autonomous driving to leverage their strong reasoning capabilities. However, extending these capabilities from 2D to full 3D understanding i…

Autonomous DrivingcounterfactualCounterfactual ReasoningDecision Making

How Many Visual Levers Drive Urban Perception? Interventional Counterfactuals via Multiple Localised Edits

2026-04-23 · Jason Tang, Stephen Law arxiv

Street-view perception models predict subjective attributes such as safety at scale, but remain correlational: they do not identify which localized visual changes would plausibly shift human judgement for a specific scen…

Image Editing

Do Metrics for Counterfactual Explanations Align with User Perception?

2026-03-16 · Felix Liedeker, Basil Ell, Philipp Cimiano, Christoph Düsing arxiv

Explainability is widely regarded as essential for trustworthy artificial intelligence systems. However, the metrics commonly used to evaluate counterfactual explanations are algorithmic evaluation metrics that are rarel…

CoSIm: Commonsense Reasoning for Counterfactual Scene Imagination

2022-07-08 · NAACL 2022 7 · Hyounghun Kim, Abhay Zala, Mohit Bansal

As humans, we can modify our assumptions about a scene by imagining alternative objects or concepts in our minds. For example, we can easily anticipate the implications of the sun being overcast by rain clouds (e.g., the…

counterfactual