paper-with-me

홈 › Papers

Attribution Explanations for Deep Neural Networks: A Theoretical Perspective

2025-08-11 · Huiqi Deng, Hongbin Pei, Quanshi Zhang, Mengnan Du arxiv

Attribution explanation is a typical approach for explaining deep neural networks (DNNs), inferring an importance or contribution score for each input variable to the final output. In recent years, numerous attribution methods have been developed to explain DNNs. However, a persistent concern remains unresolved, i.e., whether and which attribution methods faithfully reflect the actual contribution of input variables to the decision-making process. The faithfulness issue undermines the reliability and practical utility of attribution explanations. We argue that these concerns stem from three core challenges. First, difficulties arise in comparing attribution methods due to their unstructured heterogeneity, differences in heuristics, formulations, and implementations that lack a unified organization. Second, most methods lack solid theoretical underpinnings, with their rationales remaining absent, ambiguous, or unverified. Third, empirically evaluating faithfulness is challenging without ground truth. Recent theoretical advances provide a promising way to tackle these challenges, attracting increasing attention. We summarize these developments, with emphasis on three key directions: (i) Theoretical unification, which uncovers commonalities and differences among methods, enabling systematic comparisons; (ii) Theoretical rationale, clarifying the foundations of existing methods; (iii) Theoretical evaluation, rigorously proving whether methods satisfy faithfulness principles. Beyond a comprehensive review, we provide insights into how these studies help deepen theoretical understanding, inform method selection, and inspire new attribution methods. We conclude with a discussion of promising open problems for further work.

📄 PDF Abstract BibTeX arXiv:2508.07636

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Fair feature attribution for multi-output prediction: a Shapley-based perspective

2026-02-26 · Umberto Biccari, Alain Ibáñez de Opakua, José María Mato, Óscar Millet 외 arxiv

In this article, we provide an axiomatic characterization of feature attribution for multi-output predictors within the Shapley framework. While SHAP explanations are routinely computed independently for each output coor…

On the Connection between Game-Theoretic Feature Attributions and Counterfactual Explanations

2023-07-13 · Emanuele Albini, Shubham Sharma, Saumitra Mishra, Danial Dervovic 외

Explainable Artificial Intelligence (XAI) has received widespread interest in recent years, and two of the most popular types of explanations are feature attributions, and counterfactual explanations. These classes of ap…

counterfactualCounterfactual ExplanationExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)+1

Understanding Integrated Gradients with SmoothTaylor for Deep Neural Network Attribution

2020-04-22 · arXiv 2020 4 · Gary S. W. Goh, Sebastian Lapuschkin, Leander Weber, Wojciech Samek 외

Integrated Gradients as an attribution method for deep neural network models offers simple implementability. However, it suffers from noisiness of explanations which affects the ease of interpretability. The SmoothGrad t…

image-classificationImage ClassificationObject RecognitionSensitivity

Exploring Practitioner Perspectives On Training Data Attribution Explanations

2023-10-31 · Elisa Nguyen, Evgenii Kortukov, Jean Y. Song, Seong Joon Oh

Explainable AI (XAI) aims to provide insight into opaque model reasoning to humans and as such is an interdisciplinary field by nature. In this paper, we interviewed 10 practitioners to understand the possible usability …

Explainable Fall Detection for Elderly Monitoring via Temporally Stable SHAP in Skeleton-Based Human Activity Recognition

2026-04-14 · Mohammad Saleh, Azadeh Tabatabaei arxiv

Reliable fall detection in elderly care requires monitoring systems that are not only accurate but also capable of producing stable, interpretable explanations of motion dynamics, a requirement that existing post hoc exp…

Human Activity Recognition