paper-with-me

홈 › Papers

When LRP Diverges from Leave-One-Out in Transformers

2025-10-21 · Weiqiu You, Siqi Zeng, Yao-Hung Hubert Tsai, Makoto Yamada, Han Zhao arxiv

Leave-One-Out (LOO) provides an intuitive measure of feature importance but is computationally prohibitive. While Layer-Wise Relevance Propagation (LRP) offers a potentially efficient alternative, its axiomatic soundness in modern Transformers remains largely under-examined. In this work, we first show that the bilinear propagation rules used in recent advances of AttnLRP violate the implementation invariance axiom. We prove this analytically and confirm it empirically in linear attention layers. Second, we also revisit CP-LRP as a diagnostic baseline and find that bypassing relevance propagation through the softmax layer -- backpropagating relevance only through the value matrices -- significantly improves alignment with LOO, particularly in middle-to-late Transformer layers. Overall, our results suggest that (i) bilinear factorization sensitivity and (ii) softmax propagation error potentially jointly undermine LRP's ability to approximate LOO in Transformers.

📄 PDF Abstract BibTeX arXiv:2510.18810

Code (0)

등록된 구현이 없습니다.

Tasks

Feature Importance

Similar Papers 제목 키워드 기반

Humans and language models diverge when predicting repeating text

2023-10-10 · Aditya R. Vaidya, Javier Turek, Alexander G. Huth

Language models that are trained on the next-word prediction task have been shown to accurately model human behavior in word prediction and reading speed. In contrast with these findings, we present a scenario in which t…

In-Context Learning

Visual Atoms: Pre-training Vision Transformers with Sinusoidal Waves

2023-03-02 · CVPR 2023 1 · Sora Takashima, Ryo Hayamizu, Nakamasa Inoue, Hirokatsu Kataoka 외

Formula-driven supervised learning (FDSL) has been shown to be an effective method for pre-training vision transformers, where ExFractalDB-21k was shown to exceed the pre-training effect of ImageNet-21k. These studies al…

WavePaint: Resource-efficient Token-mixer for Self-supervised Inpainting

2023-07-01 · Pranav Jeevan, Dharshan Sampath Kumar, Amit Sethi

Image inpainting, which refers to the synthesis of missing regions in an image, can help restore occluded or degraded areas and also serve as a precursor task for self-supervision. The current state-of-the-art models for…

Image Inpainting

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization

2026-06-24 · Praneet Suresh, Jack Stanley, Sonia Joseph, Luca Scimeca 외 arxiv

Pre-trained transformers have demonstrated remarkable generalization abilities, at times extending beyond the scope of their training data. Yet, real-world deployments often face unexpected or adversarial data that diver…

SIGMA: Selective-Interleaved Generation with Multi-Attribute Tokens

2026-02-07 · Xiaoyan Zhang, Zechen Bai, Haofan Wang, Yiren Song arxiv

Recent unified models such as Bagel demonstrate that paired image-edit data can effectively align multiple visual tasks within a single diffusion transformer. However, these models remain limited to single-condition inpu…