paper-with-me

Papers

Aligning Faithful Interpretations with their Social Attribution

2020-06-01 · Alon Jacovi, Yoav Goldberg

We find that the requirement of model interpretations to be faithful is vague and incomplete. With interpretation by textual highlights as a case-study, we present several failure cases. Borrowing concepts from social science, we identify that the problem is a misalignment between the causal chain of decisions (causal attribution) and the attribution of human behavior to the interpretation (social attribution). We re-formulate faithfulness as an accurate attribution of causality to the model, and introduce the concept of aligned faithfulness: faithful causal chains that are aligned with their expected social behavior. The two steps of causal attribution and social attribution together complete the process of explaining behavior. With this formalization, we characterize various failures of misaligned faithful highlight interpretations, and propose an alternative causal chain to remedy the issues. Finally, we implement highlight explanations of the proposed causal format using contrastive explanations.

📄 PDF Abstract BibTeX arXiv:2006.01067

Code (1)

alonjacovi/aligned-highlights

Similar Papers 제목 키워드 기반

Sum-of-Parts: Faithful Attributions for Groups of Features

2023-10-25 · Weiqiu You, Helen Qu, Marco Gatti, Bhuvnesh Jain 외

Feature attributions explain machine learning predictions by assigning importance scores to input features. While faithful attributions accurately reflect feature contributions to the model's prediction, unfaithful ones …

Decision Makingscientific discovery

Faithful and Accurate Self-Attention Attribution for Message Passing Neural Networks via the Computation Tree Viewpoint

2024-06-07 · Yong-Min Shin, Siqing Li, Xin Cao, Won-Yong Shin

The self-attention mechanism has been adopted in various popular message passing neural networks (MPNNs), enabling the model to adaptively control the amount of information that flows along the edges of the underlying gr…

Graph Attention

Mutual Information Preserving Back-propagation: Learn to Invert for Faithful Attribution

2021-04-14 · Huiqi Deng, Na Zou, Weifu Chen, Guocan Feng 외

Back propagation based visualizations have been proposed to interpret deep neural networks (DNNs), some of which produce interpretations with good visual quality. However, there exist doubts about whether these intuitive…

Decision Making

On the Faithfulness of Vision Transformer Explanations

2024-04-01 · CVPR 2024 1 · Junyi Wu, Weitai Kang, Hao Tang, Yuan Hong 외

To interpret Vision Transformers, post-hoc explanations assign salience scores to input pixels, providing human-understandable heatmaps. However, whether these interpretations reflect true rationales behind the model's o…

Measuring Human Value Expression in Social Media Texts: Calibrated LLM Annotation and Encoder Transfer

2026-06-09 · Maria Milkova, Maksim Rudnev arxiv

Measuring subjective constructs in naturally occurring social media text requires annotation procedures that are theoretically grounded, empirically validated, and transferable to an encoder model for scalable prediction…