paper-with-me

홈 › Papers

Interpretable Vision Transformers in Image Classification via SVDA

2026-02-11 · Vasileios Arampatzakis, George Pavlidis, Nikolaos Mitianoudis, Nikos Papamarkos arxiv

Vision Transformers (ViTs) have achieved state-of-the-art performance in image classification, yet their attention mechanisms often remain opaque and exhibit dense, non-structured behaviors. In this work, we adapt our previously proposed SVD-Inspired Attention (SVDA) mechanism to the ViT architecture, introducing a geometrically grounded formulation that enhances interpretability, sparsity, and spectral structure. We apply the use of interpretability indicators -- originally proposed with SVDA -- to monitor attention dynamics during training and assess structural properties of the learned representations. Experimental evaluations on four widely used benchmarks -- CIFAR-10, FashionMNIST, CIFAR-100, and ImageNet-100 -- demonstrate that SVDA consistently yields more interpretable attention patterns without sacrificing classification accuracy. While the current framework offers descriptive insights rather than prescriptive guidance, our results establish SVDA as a comprehensive and informative tool for analyzing and developing structured attention models in computer vision. This work lays the foundation for future advances in explainable AI, spectral diagnostics, and attention-based model compression.

📄 PDF Abstract BibTeX arXiv:2602.10994

Code (0)

등록된 구현이 없습니다.

Tasks

Image ClassificationModel Compression

Similar Papers 제목 키워드 기반

Interpretable Vision Transformers in Monocular Depth Estimation via SVDA

2026-02-11 · Vasileios Arampatzakis, George Pavlidis, Nikolaos Mitianoudis, Nikos Papamarkos arxiv

Monocular depth estimation is a central problem in computer vision with applications in robotics, AR, and autonomous driving, yet the self-attention mechanisms that drive modern Transformer architectures remain opaque. W…

Monocular Depth EstimationAutonomous Driving

Multi-Source Video Domain Adaptation with Temporal Attentive Moment Alignment

2021-09-21 · Yuecong Xu, Jianfei Yang, Haozhi Cao, Keyu Wu 외

Multi-Source Domain Adaptation (MSDA) is a more practical domain adaptation scenario in real-world scenarios. It relaxes the assumption in conventional Unsupervised Domain Adaptation (UDA) that source data are sampled fr…

Domain AdaptationUnsupervised Domain Adaptation

Augmenting and Aligning Snippets for Few-Shot Video Domain Adaptation

2023-03-18 · ICCV 2023 1 · Yuecong Xu, Jianfei Yang, Yunjiao Zhou, Zhenghua Chen 외

For video models to be transferred and applied seamlessly across video tasks in varied environments, Video Unsupervised Domain Adaptation (VUDA) has been introduced to improve the robustness and transferability of video …

Action RecognitionDomain AdaptationUnsupervised Domain Adaptation

Hierarchical Vision Transformer with Prototypes for Interpretable Medical Image Classification

2025-02-13 · Luisa Gallée, Catharina Silvia Lisson, Meinrad Beer, Michael Götz

Explainability is a highly demanded requirement for applications in high-risk areas such as medicine. Vision Transformers have mainly been limited to attention extraction to provide insight into the model's reasoning. Ou…

image-classificationImage ClassificationLesion ClassificationMedical Image Classification+1

ComFe: Interpretable Image Classifiers With Foundation Models, Transformers and Component Features

2024-03-07 · Evelyn Mannix, Howard Bondell

Interpretable computer vision models are able to explain their reasoning through comparing the distances between the image patch embeddings and prototypes within a latent space. However, many of these approaches introduc…

Decoderimage-classificationImage Classificationobject-detection+1