paper-with-me

홈 › Papers

Adversarial Attacks on VQA-NLE: Exposing and Alleviating Inconsistencies in Visual Question Answering Explanations

2025-08-17 · Yahsin Yeh, Yilun Wu, Bokai Ruan, Honghan Shuai arxiv

Natural language explanations in visual question answering (VQA-NLE) aim to make black-box models more transparent by elucidating their decision-making processes. However, we find that existing VQA-NLE systems can produce inconsistent explanations and reach conclusions without genuinely understanding the underlying context, exposing weaknesses in either their inference pipeline or explanation-generation mechanism. To highlight these vulnerabilities, we not only leverage an existing adversarial strategy to perturb questions but also propose a novel strategy that minimally alters images to induce contradictory or spurious outputs. We further introduce a mitigation method that leverages external knowledge to alleviate these inconsistencies, thereby bolstering model robustness. Extensive evaluations on two standard benchmarks and two widely used VQA-NLE models underscore the effectiveness of our attacks and the potential of knowledge-based defenses, ultimately revealing pressing security and reliability concerns in current VQA-NLE systems.

📄 PDF Abstract BibTeX arXiv:2508.12430

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

KNOW How to Make Up Your Mind! Adversarially Detecting and Alleviating Inconsistencies in Natural Language Explanations

2023-06-05 · Myeongjun Jang, Bodhisattwa Prasad Majumder, Julian McAuley, Thomas Lukasiewicz 외

While recent works have been considerably improving the quality of the natural language explanations (NLEs) generated by a model to justify its predictions, there is very limited research in detecting and alleviating inc…

Adversarial Attack

Two Souls in an Adversarial Image: Towards Universal Adversarial Example Detection using Multi-view Inconsistency

2021-09-25 · Sohaib Kiani, Sana Awan, Chao Lan, Fengjun Li 외

In the evasion attacks against deep neural networks (DNN), the attacker generates adversarial instances that are visually indistinguishable from benign samples and sends them to the target DNN to trigger misclassificatio…

Adversarial Attack DetectionAdversarial DefenseAdversarial RobustnessCW Attack Detection

On Adversarial Attacks In Acoustic Drone Localization

2025-02-27 · Tamir Shor, Chaim Baskin, Alex Bronstein arxiv

Multi-rotor aerial autonomous vehicles (MAVs, more widely known as "drones") have been generating increased interest in recent years due to their growing applicability in a vast and diverse range of fields (e.g., agricul…

Visual Adversarial Attack on Vision-Language Models for Autonomous Driving

2024-11-27 · Tianyuan Zhang, Lu Wang, Xinwei Zhang, Yitong Zhang 외

Vision-language models (VLMs) have significantly advanced autonomous driving (AD) by enhancing reasoning capabilities. However, these models remain highly vulnerable to adversarial attacks. While existing research has pr…

Adversarial AttackAutonomous DrivingLanguage ModellingLarge Language Model

Demiguise Attack: Crafting Invisible Semantic Adversarial Perturbations with Perceptual Similarity

2021-07-03 · Yajie Wang, Shangbo Wu, Wenyi Jiang, Shengang Hao 외

Deep neural networks (DNNs) have been found to be vulnerable to adversarial examples. Adversarial examples are malicious images with visually imperceptible perturbations. While these carefully crafted perturbations restr…