paper-with-me

홈 › Papers

VISTA: Auditing Semantic Divergence in Vision-Language Models

2026-07-03 · Junchi Liao, Jiawen Deng, Fuji Ren arxiv

Vision-language models can exhibit visual concept-conditioned divergence: given images containing demographic features, corporate logos, or ideological symbols, some models produce unusually uniform responses that differ from what peer models say about the same input. These behaviors evade text-only audits because visual concepts cannot be isolated or substituted the way text tokens can. We present VISTA (Visual Inconsistency Screening Through Analysis), a black-box cross-model audit that couples semantic entropy with distribution-based divergence to flag model-specific anomalies. In a controlled study, we implant concept-conditioned stances in three VLMs via fine-tuning on small biased datasets and confirm that VISTA detects them. Auditing six VLMs across 19 topics, VISTA surfaces 142 high-suspicion cases (1.2%) and identifies selective refusal as a previously unreported divergence pattern, where models refuse demographic queries at rates varying from 0 to 65% across groups.

📄 PDF Abstract BibTeX arXiv:2607.02995

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MedVista3D: Vision-Language Modeling for Reducing Diagnostic Errors in 3D CT Disease Detection, Understanding and Reporting

2025-09-04 · Yuheng Li, Yenho Chen, Yuxiang Lai, Jike Zhong 외 arxiv

Radiologic diagnostic errors-under-reading errors, inattentional blindness, and communication failures-remain prevalent in clinical practice. These issues often stem from missed localized abnormalities, limited global co…

Visual Question AnsweringRepresentation LearningSpatial Reasoning

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation

2026-07-14 · Mohan Liu, Zhihao Gu, Xuanyu Chen, Haitian Zhang 외 arxiv

Vision-Language-Action (VLA) models have emerged as a powerful end-to-end paradigm for robotic manipulation by mapping language instructions and 2D visual inputs directly to actions. However, these models lack an explici…

Point Clouds

VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?

2026-02-04 · Qing'an Liu, Juntong Feng, Yuhao Wang, Xinzhe Han 외 arxiv

Vision-Language Models (VLMs) have achieved impressive performance in cross-modal understanding across textual and visual inputs, yet existing benchmarks predominantly focus on pure-text queries. In real-world scenarios,…

VISTA: A Controllable Platform for Generating and Auditing Egocentric Assistance Scenarios

2026-05-11 · Yu-Hsiang Liu, Yu-Chien Tang, An-Zi Yen arxiv

Evaluating whether AI agents can proactively assist humans in daily activities, ranging from routine household tasks to urgent safety-critical situations, requires diverse visual data. However, collecting such scenarios …

VISTA: Verifier-Informed Student-to-Teacher Adaptation for On-Policy Self-Distillation

2026-08-28 · Zewen Ding, Zezhong Wu, Zhou Tao, Shida Wang 외 arxiv

On-policy self-distillation (OPSD) improves reasoning by training a problem-only student on its own rollouts using dense token-level supervision from a privileged teacher that also sees a reference solution. However, sta…