paper-with-me

Papers

Evaluating Attribution Methods using White-Box LSTMs

2020-10-16 · EMNLP (BlackboxNLP) 2020 11 · Yiding Hao

Interpretability methods for neural networks are difficult to evaluate because we do not understand the black-box models typically used to test them. This paper proposes a framework in which interpretability methods are evaluated using manually constructed networks, which we call white-box networks, whose behavior is understood a priori. We evaluate five methods for producing attribution heatmaps by applying them to white-box LSTM classifiers for tasks based on formal languages. Although our white-box classifiers solve their tasks perfectly and transparently, we find that all five attribution methods fail to produce the expected model explanations.

📄 PDF Abstract BibTeX arXiv:2010.08606

Code (1)

yidinghao/whitebox-lstm 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

Interpretability 설명 없음
Tanh Activation 설명 없음
Sigmoid Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

The effect of whitening on explanation performance

2026-02-09 · Benedict Clark, Stoyan Karastoyanov, Rick Wilming, Stefan Haufe arxiv

Explainable Artificial Intelligence (XAI) aims to provide transparent insights into machine learning models, yet the reliability of many feature attribution methods remains a critical challenge. Prior research (Haufe et …

Saliency strikes back: How filtering out high frequencies improves white-box explanations

2023-07-18 · Sabine Muzellec, Thomas Fel, Victor Boutin, Léo Andéol 외

Attribution methods correspond to a class of explainability methods (XAI) that aim to assess how individual inputs contribute to a model's decision-making process. We have identified a significant limitation in one type …

Computational EfficiencyDecision Making

Rethinking Robustness: A New Approach to Evaluating Feature Attribution Methods

2025-12-07 · Panagiota Kiourti, Anu Singh, Preeti Duraipandian, Weichao Zhou 외 arxiv

This paper studies the robustness of feature attribution methods for deep neural networks. It challenges the current notion of attributional robustness that largely ignores the difference in the model's outputs and intro…

LSTM+Geo with xgBoost Filtering: A Novel Approach for Race and Ethnicity Imputation with Reduced Bias

2025-04-30 · S. Chalavadi, A. Pastor, T. Leitch

Accurate imputation of race and ethnicity (R&E) is crucial for analyzing disparities and informing policy. Methods like Bayesian Improved Surname Geocoding (BISG) are widely used but exhibit limitations, including system…

Imputation

Interpretable Deep Learning Model for Online Multi-touch Attribution

2020-03-26 · Dongdong Yang, Kevin Dyer, Senzhang Wang

In online advertising, users may be exposed to a range of different advertising campaigns, such as natural search or referral or organic search, before leading to a final transaction. Estimating the contribution of adver…

Deep LearningMarketing