paper-with-me

Papers

An Attention Matrix for Every Decision: Faithfulness-based Arbitration Among Multiple Attention-Based Interpretations of Transformers in Text Classification

2022-09-22 · Nikolaos Mylonas, Ioannis Mollas, Grigorios Tsoumakas

Transformers are widely used in natural language processing, where they consistently achieve state-of-the-art performance. This is mainly due to their attention-based architecture, which allows them to model rich linguistic relations between (sub)words. However, transformers are difficult to interpret. Being able to provide reasoning for its decisions is an important property for a model in domains where human lives are affected. With transformers finding wide use in such fields, the need for interpretability techniques tailored to them arises. We propose a new technique that selects the most faithful attention-based interpretation among the several ones that can be obtained by combining different head, layer and matrix operations. In addition, two variations are introduced towards (i) reducing the computational complexity, thus being faster and friendlier to the environment, and (ii) enhancing the performance in multi-label data. We further propose a new faithfulness metric that is more suitable for transformer models and exhibits high correlation with the area under the precision-recall curve based on ground truth rationales. We validate the utility of our contributions with a series of quantitative and qualitative experiments on seven datasets.

📄 PDF Abstract BibTeX arXiv:2209.10876

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationFeature ImportanceHate Speech Detectiontext-classificationText Classification

Similar Papers 제목 키워드 기반

Beyond Text Following: Repairable Arbitration Reversals in Audio-Language Models

2026-06-03 · Yichen Gao, Yiqun Zhang, Zijing Wang, Yujia Li 외 arxiv

Audio-language models (ALMs) often follow text that conflicts with audio, even when the audio evidence is clear. This raises a basic question: is the audio-supported answer unavailable, or is it represented but overridde…

Instruction Anchor: Dissecting the Mechanistic Dynamics of Modality Arbitration

2026-02-03 · Yu Zhang, Mufan Xu, Xuefeng Bai, Kehai Chen 외 arxiv

Modality following is the ability to selectively leverage multimodal contexts based on user instructions. It is fundamental to the safety and reliability of multimodal large language models (MLLMs) in real-world deployme…

End-to-end Alexa Device Arbitration

2021-12-08 · Jarred Barber, Yifeng Fan, Tao Zhang

We introduce a variant of the speaker localization problem, which we call device arbitration. In the device arbitration problem, a user utters a keyword that is detected by multiple distributed microphone arrays (smart h…

Evaluating the Faithfulness of Saliency-based Explanations for Deep Learning Models for Temporal Colour Constancy

2022-11-15 · Matteo Rizzo, Cristina Conati, Daesik Jang, Hui Hu

The opacity of deep learning models constrains their debugging and improvement. Augmenting deep models with saliency-based strategies, such as attention, has been claimed to help get a better understanding of the decisio…

Decision Making

An Arbitration Control for an Ensemble of Diversified DQN variants in Continual Reinforcement Learning

2025-09-05 · Wonseo Jang, Dongjae Kim arxiv

Deep reinforcement learning (RL) models, despite their efficiency in learning an optimal policy in static environments, easily loses previously learned knowledge (i.e., catastrophic forgetting). It leads RL models to poo…

Reinforcement Learning