paper-with-me

홈 › Papers

Interpretability is a Kind of Safety: An Interpreter-based Ensemble for Adversary Defense

2023-04-14 · Jingyuan Wang, Yufan Wu, Mingxuan Li, Xin Lin, Junjie Wu, Chao Li

While having achieved great success in rich real-life applications, deep neural network (DNN) models have long been criticized for their vulnerability to adversarial attacks. Tremendous research efforts have been dedicated to mitigating the threats of adversarial attacks, but the essential trait of adversarial examples is not yet clear, and most existing methods are yet vulnerable to hybrid attacks and suffer from counterattacks. In light of this, in this paper, we first reveal a gradient-based correlation between sensitivity analysis-based DNN interpreters and the generation process of adversarial examples, which indicates the Achilles's heel of adversarial attacks and sheds light on linking together the two long-standing challenges of DNN: fragility and unexplainability. We then propose an interpreter-based ensemble framework called X-Ensemble for robust adversary defense. X-Ensemble adopts a novel detection-rectification process and features in building multiple sub-detectors and a rectifier upon various types of interpretation information toward target classifiers. Moreover, X-Ensemble employs the Random Forests (RF) model to combine sub-detectors into an ensemble detector for adversarial hybrid attacks defense. The non-differentiable property of RF further makes it a precious choice against the counterattack of adversaries. Extensive experiments under various types of state-of-the-art attacks and diverse attack scenarios demonstrate the advantages of X-Ensemble to competitive baseline methods.

📄 PDF Abstract BibTeX arXiv:2304.06919

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Hybrid CNN -Interpreter: Interpret local and global contexts for CNN-based Models

2022-10-31 · Wenli Yang, Guan Huang, Renjie Li, Jiahao Yu 외

Convolutional neural network (CNN) models have seen advanced improvements in performance in various domains, but lack of interpretability is a major barrier to assurance and regulation during operation for acceptance and…

Feature Correlation

A Framework to Learn with Interpretation

2020-10-19 · NeurIPS 2021 12 · Jayneel Parekh, Pavlo Mozharovskyi, Florence d'Alché-Buc

To tackle interpretability in deep learning, we present a novel framework to jointly learn a predictive model and its associated interpretation model. The interpreter provides both local and global interpretability about…

AttributeDecision Making

Grounded Vision-Language Interpreter for Integrated Task and Motion Planning

2025-06-03 · Jeremy Siburian, Keisuke Shirai, Cristian C. Beltran-Hernandez, Masashi Hamaya 외

While recent advances in vision-language models (VLMs) have accelerated the development of language-guided robot planners, their black-box nature often lacks safety guarantees and interpretability crucial for real-world …

Motion PlanningTask and Motion PlanningTask Planning

Radical AI Interpretability

2026-06-25 · Daniel A. Herrmann, Benjamin A. Levinstein arxiv

We develop a framework for interpreting AI systems as agents, drawing on the philosophical tradition of radical interpretation and the tools of mechanistic interpretability. The core question is: given the computational …

DISentangled Counterfactual Visual interpretER (DISCOVER) generalizes to natural images

2024-06-22 · Oded Rotem, Assaf Zaritsky

We recently presented DISentangled COunterfactual Visual interpretER (DISCOVER), a method toward systematic visual interpretability of image-based classification models and demonstrated its applicability to two biomedica…

counterfactual