paper-with-me

홈 › Papers

Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time

2026-09-04 · Lisa Bouger, Yannick Teglia, Philippe Loubet Moundi arxiv

We propose an influence score to quantify the contribution of attention heads to classification decisions in Transformer-based models designed for prompt injection detection. The score combines directional influence on the logits with structural contribution within the residual stream, enabling a multi-scale analysis at the head, layer, and network levels. Applied to a DeBERTa model specialized for prompt injection detection, our framework reveals distinct decision behaviours between correct and erroneous predictions. Our method provides an effective compromise between fine-grained circuit analysis and global output-based methods, and offers a systematic way to study decision mechanisms in Transformer classifiers.

📄 PDF Abstract BibTeX arXiv:2609.05074

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Step-resolved data attribution for looped transformers

2026-02-10 · Georgios Kaissis, David Mildenberger, Juan Felipe Gomez, Martin J. Menten 외 arxiv

We study how individual training examples shape the internal computation of looped transformers, where a shared block is applied for $τ$ recurrent iterations to enable latent reasoning. Existing training-data influence e…

Disentangling Visual Transformers: Patch-level Interpretability for Image Classification

2025-02-24 · Guillaume Jeanneret, Loïc Simon, Frédéric Jurie

Visual transformers have achieved remarkable performance in image classification tasks, but this performance gain has come at the cost of interpretability. One of the main obstacles to the interpretation of transformers …

image-classificationImage Classification

Attributions All the Way Down? The Metagame of Interpretability

2026-05-07 · Hubert Baniecki, Przemyslaw Biecek, Fabian Fumagalli arxiv

We introduce the metagame, a conceptual framework for quantifying second-order interaction effects of model explanations. For any first-order attribution $φ(f)$ explaining a model $f$, we measure the directional influenc…

Evaluating self-attention interpretability through human-grounded experimental protocol

2023-03-27 · Milan Bhan, Nina Achache, Victor Legrand, Annabelle Blangero 외

Attention mechanisms have played a crucial role in the development of complex architectures such as Transformers in natural language processing. However, Transformers remain hard to interpret and are considered as black-…

X-VMamba: Explainable Vision Mamba

2025-11-16 · Mohamed A. Mabrok, Yalda Zafari arxiv

State Space Models (SSMs), particularly the Mamba architecture, have recently emerged as powerful alternatives to Transformers for sequence modeling, offering linear computational complexity while achieving competitive p…