paper-with-me

Papers

[Re] Improving Interpretation Faithfulness for Vision Transformers

2025-09-18 · Izabela Kurek, Wojciech Trejter, Stipe Frkovic, Andro Erdelez arxiv

This work aims to reproduce the results of Faithful Vision Transformers (FViTs) proposed by arXiv:2311.17983 alongside interpretability methods for Vision Transformers from arXiv:2012.09838 and Xu (2022) et al. We investigate claims made by arXiv:2311.17983, namely that the usage of Diffusion Denoised Smoothing (DDS) improves interpretability robustness to (1) attacks in a segmentation task and (2) perturbation and attacks in a classification task. We also extend the original study by investigating the authors' claims that adding DDS to any interpretability method can improve its robustness under attack. This is tested on baseline methods and the recently proposed Attribution Rollout method. In addition, we measure the computational costs and environmental impact of obtaining an FViT through DDS. Our results broadly agree with the original study's findings, although minor discrepancies were found and discussed.

📄 PDF Abstract BibTeX arXiv:2509.14846

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On the Faithfulness of Vision Transformer Explanations

2024-04-01 · CVPR 2024 1 · Junyi Wu, Weitai Kang, Hao Tang, Yuan Hong 외

To interpret Vision Transformers, post-hoc explanations assign salience scores to input pixels, providing human-understandable heatmaps. However, whether these interpretations reflect true rationales behind the model's o…

VISION DIFFMASK: Faithful Interpretation of Vision Transformers with Differentiable Patch Masking

2023-04-13 · Angelos Nalmpantis, Apostolos Panagiotopoulos, John Gkountouras, Konstantinos Papakostas 외

The lack of interpretability of the Vision Transformer may hinder its use in critical real-world applications despite its effectiveness. To overcome this issue, we propose a post-hoc interpretability method called VISION…

Improving Interpretation Faithfulness for Vision Transformers

2023-11-29 · Lijie Hu, Yixin Liu, Ninghao Liu, Mengdi Huai 외

Vision Transformers (ViTs) have achieved state-of-the-art performance for various vision tasks. One reason behind the success lies in their ability to provide plausible innate explanations for the behavior of neural arch…

Denoising

An Attention Matrix for Every Decision: Faithfulness-based Arbitration Among Multiple Attention-Based Interpretations of Transformers in Text Classification

2022-09-22 · Nikolaos Mylonas, Ioannis Mollas, Grigorios Tsoumakas

Transformers are widely used in natural language processing, where they consistently achieve state-of-the-art performance. This is mainly due to their attention-based architecture, which allows them to model rich linguis…

ClassificationFeature ImportanceHate Speech Detectiontext-classification+1

Towards Evaluating Explanations of Vision Transformers for Medical Imaging

2023-04-12 · Piotr Komorowski, Hubert Baniecki, Przemysław Biecek

As deep learning models increasingly find applications in critical domains such as medical imaging, the need for transparent and trustworthy decision-making becomes paramount. Many explainability methods provide insights…

Decision Makingimage-classificationImage Classification