paper-with-me

Papers

Adversarial Attack for Explanation Robustness of Rationalization Models

2024-08-20 · Yuankai Zhang, Lingxiao Kong, Haozhao Wang, Ruixuan Li, Jun Wang, Yuhua Li, Wei Liu

Rationalization models, which select a subset of input text as rationale-crucial for humans to understand and trust predictions-have recently emerged as a prominent research area in eXplainable Artificial Intelligence. However, most of previous studies mainly focus on improving the quality of the rationale, ignoring its robustness to malicious attack. Specifically, whether the rationalization models can still generate high-quality rationale under the adversarial attack remains unknown. To explore this, this paper proposes UAT2E, which aims to undermine the explainability of rationalization models without altering their predictions, thereby eliciting distrust in these models from human users. UAT2E employs the gradient-based search on triggers and then inserts them into the original input to conduct both the non-target and target attack. Experimental results on five datasets reveal the vulnerability of rationalization models in terms of explanation, where they tend to select more meaningless tokens under attacks. Based on this, we make a series of recommendations for improving rationalization models in terms of explanation.

📄 PDF Abstract BibTeX arXiv:2408.10795

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial AttackExplainable artificial intelligence

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Can Rationalization Improve Robustness?

2022-04-25 · NAACL 2022 7 · Howard Chen, Jacqueline He, Karthik Narasimhan, Danqi Chen

A growing line of work has investigated the development of neural NLP models that can produce rationales--subsets of input that can explain their model predictions. In this paper, we ask whether such rationale models can…

Sentence

Can Rationalization Improve Robustness?

2021-11-16 · ACL ARR November 2021 11 · Anonymous

A growing line of work has investigated the development of neural NLP models that can produce rationales---subsets of input that can explain their model predictions. In this paper, we ask whether such rationale models ca…

Sentence

Overcoming Adversarial Attacks for Human-in-the-Loop Applications

2023-06-09 · Ryan McCoppin, Marla Kennedy, Platon Lukyanenko, Sean Kennedy

Including human analysis has the potential to positively affect the robustness of Deep Neural Networks and is relatively unexplored in the Adversarial Machine Learning literature. Neural network visual explanation maps h…

JECA^2: Judgment-Explanation Consistent Adversarial Attack against Forensic Vision-Language Models

2026-05-27 · Jiachen Qian arxiv

Forensic vision-language models (VLMs) have recently been developed to detect image tampering and provide natural-language explanations. However, their robustness against adversarial manipulation remains underexplored. E…

Adversarial Attack

Rationalization: A Neural Machine Translation Approach to Generating Natural Language Explanations

2017-02-25 · Upol Ehsan, Brent Harrison, Larry Chan, Mark O. Riedl

We introduce AI rationalization, an approach for generating explanations of autonomous system behavior as if a human had performed the behavior. We describe a rationalization technique that uses neural machine translatio…

Explanation GenerationMachine TranslationTranslation