paper-with-me

홈 › Papers

From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA

2026-06-29 · Sena Korkut, María Alejandra Bravo Sarmiento, Sanghwan Kim, Zeynep Akata arxiv

High benchmark accuracy does not guarantee genuine use of visual evidence. We study this problem in traffic accident Video Question Answering (VideoQA), where correct answers should depend on scene-specific visual evidence but may instead be inferred from textual shortcuts. Through an audit of four public benchmarks, we find that several recent open-weight Vision-Language Models (VLMs) perform competitively, and sometimes better, without video input. On the MM-AU benchmark, removing video consistently improves accuracy, and adding more frames further degrades performance. To quantify visual dependence, we introduce two dataset-level diagnostics: Blind Gap, measuring above-chance text-only performance, and Visual Gain, measuring the marginal benefit of adding video. We further propose an instance-level Shortcut Score that combines text-only confidence with visual necessity signals, enabling continuous, training-free filtering of shortcut-prone questions. The resulting subsets reduce shortcut bias and improve visual grounding. Our findings reveal large differences in grounding quality across benchmarks and show that visually grounded evaluation, not just high accuracy, is essential in safety-critical VideoQA.

📄 PDF Abstract BibTeX arXiv:2606.30220

Code (0)

등록된 구현이 없습니다.

Tasks

Video Question AnsweringVisual Grounding

Similar Papers 제목 키워드 기반

Evian: Towards Explainable Visual Instruction-tuning Data Auditing

2026-04-22 · Zimu Jia, Mingjie Xu, Andrew Estornell, Jiaheng Wei arxiv

The efficacy of Large Vision-Language Models (LVLMs) is critically dependent on the quality of their training data, requiring a precise balance between visual fidelity and instruction-following capability. Existing datas…

Logical Fallacies

Video-KTR: Reinforcing Video Reasoning via Key Token Attribution

2026-01-27 · Ziyue Wang, Sheng Jin, Zhongrong Zuo, Jiawei Wu 외 arxiv

Reinforcement learning (RL) has shown strong potential for enhancing reasoning in multimodal large language models, yet existing video reasoning methods often rely on coarse sequence-level rewards or single-factor token …

Reinforcement Learning

A Real-Time Cross-modality Correlation Filtering Method for Referring Expression Comprehension

2019-09-16 · CVPR 2020 6 · Yue Liao, Si Liu, Guanbin Li, Fei Wang 외

Referring expression comprehension aims to localize the object instance described by a natural language expression. Current referring expression methods have achieved good performance. However, none of them is able to ac…

Referring ExpressionReferring Expression Comprehension

FF-LOGO: Cross-Modality Point Cloud Registration with Feature Filtering and Local to Global Optimization

2023-09-16 · Nan Ma, Mohan Wang, Yiheng Han, Yong-Jin Liu

Cross-modality point cloud registration is confronted with significant challenges due to inherent differences in modalities between different sensors. We propose a cross-modality point cloud registration framework FF-LOG…

Feature Correlationglobal-optimizationPoint Cloud Registration

Who Gets Flagged? The Pluralistic Evaluation Gap in AI Content Watermarking

2026-04-15 · Alexander Nemecek, Osama Zafar, Yuqiao Xu, Wenbiao Li 외 arxiv

Watermarking is becoming the default mechanism for AI content authentication, with governance policies and frameworks referencing it as infrastructure for content provenance. Yet across text, image, and audio modalities,…