paper-with-me

홈 › Papers

Improving the Faithfulness of Attention-based Explanations with Task-specific Information for Text Classification

2021-05-06 · ACL 2021 5 · George Chrysostomou, Nikolaos Aletras

Neural network architectures in natural language processing often use attention mechanisms to produce probability distributions over input token representations. Attention has empirically been demonstrated to improve performance in various tasks, while its weights have been extensively used as explanations for model predictions. Recent studies (Jain and Wallace, 2019; Serrano and Smith, 2019; Wiegreffe and Pinter, 2019) have showed that it cannot generally be considered as a faithful explanation (Jacovi and Goldberg, 2020) across encoders and tasks. In this paper, we seek to improve the faithfulness of attention-based explanations for text classification. We achieve this by proposing a new family of Task-Scaling (TaSc) mechanisms that learn task-specific non-contextualised information to scale the original attention weights. Evaluation tests for explanation faithfulness, show that the three proposed variants of TaSc improve attention-based explanations across two attention mechanisms, five encoders and five text classification datasets without sacrificing predictive performance. Finally, we demonstrate that TaSc consistently provides more faithful attention-based explanations compared to three widely-used interpretability techniques.

📄 PDF Abstract BibTeX arXiv:2105.02657

Code (1)

gchrysostomou/tasc pytorch

Tasks

text-classificationText Classification

Similar Papers 제목 키워드 기반

On the Faithfulness of Vision Transformer Explanations

2024-04-01 · CVPR 2024 1 · Junyi Wu, Weitai Kang, Hao Tang, Yuan Hong 외

To interpret Vision Transformers, post-hoc explanations assign salience scores to input pixels, providing human-understandable heatmaps. However, whether these interpretations reflect true rationales behind the model's o…

Human Attention-Guided Explainable Artificial Intelligence for Computer Vision Models

2023-05-05 · Guoyang Liu, Jindi Zhang, Antoni B. Chan, Janet H. Hsiao

We examined whether embedding human attention knowledge into saliency-based explainable AI (XAI) methods for computer vision models could enhance their plausibility and faithfulness. We first developed new gradient-based…

ClassificationExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)image-classification+4

Enjoy the Salience: Towards Better Transformer-based Faithful Explanations with Word Salience

2021-08-31 · EMNLP 2021 11 · George Chrysostomou, Nikolaos Aletras

Pretrained transformer-based models such as BERT have demonstrated state-of-the-art predictive performance when adapted into a range of natural language processing tasks. An open problem is how to improve the faithfulnes…

Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations

2026-02-13 · Ajay Pravin Mahale arxiv

Mechanistic interpretability identifies internal circuits responsible for model behaviors, yet translating these findings into human-understandable explanations remains an open problem. We present a pipeline that bridges…

Evaluating Human Alignment and Model Faithfulness of LLM Rationale

2024-06-28 · Mohsen Fayyaz, Fan Yin, Jiao Sun, Nanyun Peng

We study how well large language models (LLMs) explain their generations through rationales -- a set of tokens extracted from the input text that reflect the decision-making process of LLMs. Specifically, we systematical…

Decision Making