paper-with-me

홈 › Papers

On the Robustness of Removal-Based Feature Attributions

2023-06-12 · NeurIPS 2023 11

To explain predictions made by complex machine learning models, many feature attribution methods have been developed that assign importance scores to input features. Some recent work challenges the robustness of these methods by showing that they are sensitive to input and model perturbations, while other work addresses this issue by proposing robust attribution methods. However, previous work on attribution robustness has focused primarily on gradient-based feature attributions, whereas the robustness of removal-based attribution methods is not currently well understood. To bridge this gap, we theoretically characterize the robustness properties of removal-based feature attributions. Specifically, we provide a unified analysis of such methods and derive upper bounds for the difference between intact and perturbed attributions, under settings of both input and model perturbations. Our empirical results on synthetic and real-world data validate our theoretical results and demonstrate their practical implications, including the ability to increase attribution robustness by improving the model's Lipschitz regularity.

📄 PDF Abstract BibTeX arXiv:2306.07462

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Algorithms to estimate Shapley value feature attributions

2022-07-15 · Hugh Chen, Ian C. Covert, Scott M. Lundberg, Su-In Lee

Feature attributions based on the Shapley value are popular for explaining machine learning models; however, their estimation is complex from both a theoretical and computational standpoint. We disentangle this complexit…

Rethinking Robustness of Model Attributions

2023-12-16 · Sandesh Kamath, Sankalp Mittal, Amit Deshpande, Vineeth N Balasubramanian

For machine learning models to be reliable and trustworthy, their decisions must be interpretable. As these models find increasing use in safety-critical applications, it is important that not just the model predictions …

Diversitymodel

Provably Better Explanations with Optimized Aggregation of Feature Attributions

2024-06-07 · Thomas Decker, Ananta R. Bhattarai, Jindong Gu, Volker Tresp 외

Using feature attributions for post-hoc explanations is a common practice to understand and verify the predictions of opaque machine learning models. Despite the numerous techniques available, individual methods often pr…

Explainable Learning with Gaussian Processes

2024-03-11 · Kurt Butler, Guanchao Feng, Petar M. Djuric

The field of explainable artificial intelligence (XAI) attempts to develop methods that provide insight into how complicated machine learning methods make predictions. Many methods of explanation have focused on the conc…

Explainable artificial intelligenceExplainable Artificial Intelligence (XAI)Gaussian ProcessesGPR

Certified $\ell_2$ Attribution Robustness via Uniformly Smoothed Attributions

2024-05-10 · Fan Wang, Adams Wai-Kin Kong

Model attribution is a popular tool to explain the rationales behind model predictions. However, recent work suggests that the attributions are vulnerable to minute perturbations, which can be added to input samples to f…