paper-with-me

홈 › Papers

Explainability of Speech Recognition Transformers via Gradient-based Attention Visualization

2023-06-02 · IEEE Transactions on Multimedia 2023 6 · Tianli Sun, Haonan Chen, Guosheng Hu, Lianghua He, Cairong Zhao

In vision Transformers, attention visualization methods are used to generate heatmaps highlighting the class-corresponding areas in input images, which offers explanations on how the models make predictions. However, it is not so applicable for explaining automatic speech recognition (ASR) Transformers. An ASR Transformer makes a particular prediction for every input token to form a sentence, but a vision Transformer only makes an overall classification for the input data. Therefore, traditional attention visualization methods may fail in ASR Transformers. In this work, we propose a novel attention visualization method in ASR Transformers and try to explain which frames of the audio result in the output text. Inspired by the model explainability, we also explore ways of improving the effectiveness of the ASR model. Comparing with other Transformer attention visualization methods, our method is more efficient and intuitively understandable, which unravels the attention calculation from information flow of Transformer attention modules. In addition, we demonstrate the utilization of visualization result in three ways: (1) We visualize attention with respect to connectionist temporal classification (CTC) loss to train an ASR model with adversarial attention erasing regularization, which effectively decreases the word error rate (WER) of the model and improves its generalization capability. (2) We visualize the attention on some specific words, interpreting the model by effectively demonstrating the semantic and grammar relationships between these words. (3) Similarly, we analyze how the model manage to distinguish homophones, using contrastive explanation with respect to homophones.

📄 PDF Abstract BibTeX

Code (1)

Vill-Lab/2023-TMM-Grad-SAS pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Sentencespeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Residual Connection 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

LeGrad: An Explainability Method for Vision Transformers via Feature Formation Sensitivity

2024-04-04 · Walid Bousselham, Angie Boggust, Sofian Chaybouti, Hendrik Strobelt 외

Vision Transformers (ViTs), with their ability to model long-range dependencies through self-attention mechanisms, have become a standard architecture in computer vision. However, the interpretability of these models rem…

Sensitivity

On the Usefulness of Self-Attention for Automatic Speech Recognition with Transformers

2020-11-08 · Shucong Zhang, Erfan Loweimi, Peter Bell, Steve Renals

Self-attention models such as Transformers, which can capture temporal relationships without being limited by the distance between events, have given competitive speech recognition results. However, we note the range of …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

XNOR-FORMER: Learning Accurate Approximations in Long Speech Transformers

2022-10-29 · Roshan Sharma, Bhiksha Raj

Transformers are among the state of the art for many tasks in speech, vision, and natural language processing, among others. Self-attentions, which are crucial contributors to this performance have quadratic computationa…

speech-recognitionSpeech Recognition

Evaluating the Explainability of Vision Transformers in Medical Imaging

2025-10-13 · Leili Barekatain, Ben Glocker arxiv

Understanding model decisions is crucial in medical imaging, where interpretability directly impacts clinical trust and adoption. Vision Transformers (ViTs) have demonstrated state-of-the-art performance in diagnostic im…

Image Classification

ViTmiX: Vision Transformer Explainability Augmented by Mixed Visualization Methods

2024-12-18 · Eduard Hogea, Darian M. Onchis, Ana Coporan, Adina Magda Florea 외

Recent advancements in Vision Transformers (ViT) have demonstrated exceptional results in various visual recognition tasks, owing to their ability to capture long-range dependencies in images through self-attention mecha…

Decision MakingExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)Semantic Segmentation