paper-with-me

홈 › Papers

Logic Traps in Evaluating Attribution Scores

2021-09-12 · ACL 2022 5 · Yiming Ju, Yuanzhe Zhang, Zhao Yang, Zhongtao Jiang, Kang Liu, Jun Zhao

Modern deep learning models are notoriously opaque, which has motivated the development of methods for interpreting how deep models predict. This goal is usually approached with attribution method, which assesses the influence of features on model predictions. As an explanation method, the evaluation criteria of attribution methods is how accurately it re-reflects the actual reasoning process of the model (faithfulness). Meanwhile, since the reasoning process of deep models is inaccessible, researchers design various evaluation methods to demonstrate their arguments. However, some crucial logic traps in these evaluation methods are ignored in most works, causing inaccurate evaluation and unfair comparison. This paper systematically reviews existing methods for evaluating attribution scores and summarizes the logic traps in these methods. We further conduct experiments to demonstrate the existence of each logic trap. Through both the theoretical and experimental analysis, we hope to increase attention on the inaccurate evaluation of attribution scores. Moreover, with this paper, we suggest stopping focusing on improving performance under unreliable evaluation systems and starting efforts on reducing the impact of proposed logic traps

📄 PDF Abstract BibTeX arXiv:2109.05463

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Logic Traps in Evaluating Attribution Scores

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Modern deep learning models are notoriously opaque, which has motivated the development of methods for interpreting how deep models predict. This goal is usually approached with attribution method, which assesses the inf…

Evaluating Explainable AI Attribution Methods in Neural Machine Translation via Attention-Guided Knowledge Distillation

2026-03-11 · Aria Nourbakhsh, Salima Lamsiyah, Adelaide Danilov, Christoph Schommer arxiv

The study of the attribution of input features to the output of neural network models is an active area of research. While numerous Explainable AI (XAI) techniques have been proposed to interpret these models, the system…

Knowledge DistillationMachine Translation

Attribution-based Explanations for Markov Decision Processes

2026-05-10 · Paul Kobialka, Andrea Pferscher, Francesco Leofante, Erika Ábrahám 외 arxiv

Attribution techniques explain the outcome of an AI model by assigning a numerical score to its inputs. So far, these techniques have mainly focused on attributing importance to static input features at a single point in…

Feature Attribution from First Principles

2025-05-30 · Magamed Taimeskhanov, Damien Garreau

Feature attribution methods are a popular approach to explain the behavior of machine learning models. They assign importance scores to each input feature, quantifying their influence on the model's prediction. However, …

CRiT-QA: Evaluating Multi-hop Reasoning with Counterfactual Chains and Distractor Traps

2026-07-12 · JungMin Yun, JuneHyoung Kwon, YoungBin Kim arxiv

Evaluating the multi-hop reasoning capabilities of large language models remains a significant challenge. Although current models achieve strong results on existing multi-hop question answering datasets, such performance…

Multi-hop Question Answering