paper-with-me

홈 › Papers

Towards trustworthy explanations with gradient-based attribution methods

2021-09-24 · NeurIPS Workshop AI4Scien 2021 12 · Ethan Louis Labelson, Rohit Tripathy, Peter K Koo

The low interpretability of deep neural networks (DNNs) remains a key barrier to their wide-spread adoption in the sciences. Attribution methods offer a promising solution, providing feature importance scores that serve as first-order model explanations for a given input. In practice, gradient-based attribution methods, such as saliency maps, can yield noisy importance scores depending on model architecture and training procedure. Here we explore how various regularization techniques affect model explanations with saliency maps using synthetic regulatory genomic data, which allows us to quantitatively assess the efficacy of attribution maps. Strikingly, we find that generalization performance does not imply better saliency explanations; though unlike before, we do not observe a clear tradeoff. Interestingly, we find that conventional regularization strategies, when tuned appropriately, can yield high generalization and interpretability performance, similar to what can be achieved with more sophisticated techniques, such as manifold mixup. Our work challenges the conventional knowledge that model selection should be based on test performance; another criterion is needed to sub-select models ideally suited for downstream post hoc interpretability for scientific discovery.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Feature ImportanceModel Selectionscientific discovery

Methods 이 논문이 사용한 방법론

HOC 설명 없음

Similar Papers 제목 키워드 기반

Sampling Matters in Explanations: Towards Trustworthy Attribution Analysis Building Block in Visual Models through Maximizing Explanation Certainty

2025-06-24 · Róisín Luo, James McDermott, Colm O'Riordan

Image attribution analysis seeks to highlight the feature representations learned by visual models such that the highlighted feature maps can reflect the pixel-wise importance of inputs. Gradient integration is a buildin…

Image Attribution

Towards Faithful Explanations for Text Classification with Robustness Improvement and Explanation Guided Training

2023-12-29 · Dongfang Li, Baotian Hu, Qingcai Chen, Shan He

Feature attribution methods highlight the important input tokens as explanations to model predictions, which have been widely applied to deep neural networks towards trustworthy AI. However, recent works show that explan…

text-classificationText Classification

Training for Trustworthy Saliency Maps: Adversarial Training Meets Feature-Map Smoothing

2026-03-07 · Dipkamal Bhusal, Md Tanvirul Alam, Nidhi Rastogi arxiv

Gradient-based saliency methods such as Vanilla Gradient (VG) and Integrated Gradients (IG) are widely used to explain image classifiers, yet the resulting maps are often noisy and unstable, limiting their usefulness in …

XAI-Grounded Explanation Generation for Speech Deepfake Detection with Training-Free Multimodal Large Language Models

2026-06-15 · Yupei Li, Qiyang Sun, Xiaoliang Wu, Chenxi Wang 외 arxiv

Speech deepfake detection (SDD) systems require trustworthy explanations for reliable decision-making. Existing explanation ways mainly fall into two categories. Traditional explainable AI (XAI), such as gradient-based a…

Explanation GenerationDeepFake Detection

Attributional Robustness Training using Input-Gradient Spatial Alignment

2019-11-29 · ECCV 2020 8 · Mayank Singh, Nupur Kumari, Puneet Mangla, Abhishek Sinha 외

Interpretability is an emerging area of research in trustworthy machine learning. Safe deployment of machine learning system mandates that the prediction and its explanation be reliable and robust. Recently, it has been …

BIG-bench Machine LearningObject LocalizationTripletWeakly-Supervised Object Localization