paper-with-me

홈 › Papers

Human-inspired Explanations for Vision Transformers and Convolutional Neural Networks

2024-08-04 · Mahadev Prasad Panda, Matteo Tiezzi, Martina Vilas, Gemma Roig, Bjoern M. Eskofier, Dario Zanca

We introduce Foveation-based Explanations (FovEx), a novel human-inspired visual explainability (XAI) method for Deep Neural Networks. Our method achieves state-of-the-art performance on both transformer (on 4 out of 5 metrics) and convolutional models (on 3 out of 5 metrics), demonstrating its versatility. Furthermore, we show the alignment between the explanation map produced by FovEx and human gaze patterns (+14\% in NSS compared to RISE, +203\% in NSS compared to gradCAM), enhancing our confidence in FovEx's ability to close the interpretation gap between humans and machines.

📄 PDF Abstract BibTeX arXiv:2408.02123

Code (1)

mahadev1995/FovEx 공식 구현 pytorch

Tasks

Foveation

Similar Papers 제목 키워드 기반

B-cosification: Transforming Deep Neural Networks to be Inherently Interpretable

2024-11-01 · Shreyash Arya, Sukrut Rao, Moritz Böhle, Bernt Schiele

B-cos Networks have been shown to be effective for obtaining highly human interpretable explanations of model decisions by architecturally enforcing stronger alignment between inputs and weight. B-cos variants of convolu…

Evaluating Graphical Perception Capabilities of Vision Transformers

2026-02-20 · Poonam Poonam, Pere-Pau Vázquez, Timo Ropinski arxiv

Vision Transformers, ViTs, have emerged as a powerful alternative to convolutional neural networks, CNNs, in a variety of image-based tasks. While CNNs have previously been evaluated for their ability to perform graphica…

3D Human Pose Estimation with Spatial and Temporal Transformers

2021-03-18 · ICCV 2021 10 · Ce Zheng, Sijie Zhu, Matias Mendieta, Taojiannan Yang 외

Transformer architectures have become the model of choice in natural language processing and are now being introduced into computer vision tasks such as image classification, object detection, and semantic segmentation. …

3D Human Pose Estimationimage-classificationImage ClassificationMonocular 3D Human Pose Estimation+4

T-TAME: Trainable Attention Mechanism for Explaining Convolutional Networks and Vision Transformers

2024-03-07 · Mariano V. Ntrougkas, Nikolaos Gkalelis, Vasileios Mezaris

The development and adoption of Vision Transformers and other deep-learning architectures for image classification tasks has been rapid. However, the "black box" nature of neural networks is a barrier to adoption in appl…

Explanation Generationimage-classification

ViT-P: Rethinking Data-efficient Vision Transformers from Locality

2022-03-04 · Bin Chen, Ran Wang, Di Ming, Xin Feng

Recent advances of Transformers have brought new trust to computer vision tasks. However, on small dataset, Transformers is hard to train and has lower performance than convolutional neural networks. We make vision trans…