paper-with-me

Papers

Rethinking Stability for Attribution-based Explanations

2022-03-14 · Chirag Agarwal, Nari Johnson, Martin Pawelczyk, Satyapriya Krishna, Eshika Saxena, Marinka Zitnik, Himabindu Lakkaraju

As attribution-based explanation methods are increasingly used to establish model trustworthiness in high-stakes situations, it is critical to ensure that these explanations are stable, e.g., robust to infinitesimal perturbations to an input. However, previous works have shown that state-of-the-art explanation methods generate unstable explanations. Here, we introduce metrics to quantify the stability of an explanation and show that several popular explanation methods are unstable. In particular, we propose new Relative Stability metrics that measure the change in output explanation with respect to change in input, model representation, or output of the underlying predictor. Finally, our experimental evaluation with three real-world datasets demonstrates interesting insights for seven explanation methods and different stability metrics.

📄 PDF Abstract BibTeX arXiv:2203.06877

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Rethinking Explainability in the Era of Multimodal AI

2025-06-16 · Chirag Agarwal

While multimodal AI systems (models jointly trained on heterogeneous data types such as text, time series, graphs, and images) have become ubiquitous and achieved remarkable performance across high-stakes applications, t…

Can We Really Trust Explanations? Evaluating the Stability of Feature Attribution Explanation Methods via Adversarial Attack

2022-10-01 · CCL 2022 10 · Yang Zhao, Zhang Yuanzhe, Jiang Zhongtao, Ju Yiming 외

“Explanations can increase the transparency of neural networks and make them more trustworthy. However, can we really trust explanations generated by the existing explanation methods? If the explanation methods are not s…

Adversarial Attack

Attention Mechanisms Don't Learn Additive Models: Rethinking Feature Importance for Transformers

2024-05-22 · Tobias Leemann, Alina Fastowski, Felix Pfeiffer, Gjergji Kasneci

We address the critical challenge of applying feature attribution methods to the transformer architecture, which dominates current applications in natural language processing and beyond. Traditional attribution methods t…

Additive modelsFeature Importance

Feature Attribution Stability Suite: How Stable Are Post-Hoc Attributions?

2026-04-02 · Kamalasankari Subramaniakuppusamy, Jugal Gajjar arxiv

Post-hoc feature attribution methods are widely deployed in safety-critical vision systems, yet their stability under realistic input perturbations remains poorly characterized. Existing metrics evaluate explanations pri…

Rethinking Robustness of Model Attributions

2023-12-16 · Sandesh Kamath, Sankalp Mittal, Amit Deshpande, Vineeth N Balasubramanian

For machine learning models to be reliable and trustworthy, their decisions must be interpretable. As these models find increasing use in safety-critical applications, it is important that not just the model predictions …

Diversitymodel