paper-with-me

Papers

Visualizing Automatic Speech Recognition -- Means for a Better Understanding?

2022-02-01 · Karla Markert, Romain Parracone, Mykhailo Kulakov, Philip Sperl, Ching-Yu Kao, Konstantin Böttinger

Automatic speech recognition (ASR) is improving ever more at mimicking human speech processing. The functioning of ASR, however, remains to a large extent obfuscated by the complex structure of the deep neural networks (DNNs) they are based on. In this paper, we show how so-called attribution methods, that we import from image recognition and suitably adapt to handle audio data, can help to clarify the working of ASR. Taking DeepSpeech, an end-to-end model for ASR, as a case study, we show how these techniques help to visualize which features of the input are the most influential in determining the output. We focus on three visualization techniques: Layer-wise Relevance Propagation (LRP), Saliency Maps, and Shapley Additive Explanations (SHAP). We compare these methods and discuss potential further applications, such as in the detection of adversarial examples.

📄 PDF Abstract BibTeX arXiv:2202.00673

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Gradient-Adjusted Neuron Activation Profiles for Comprehensive Introspection of Convolutional Speech Recognition Models

2020-02-19 · Andreas Krug, Sebastian Stober

Deep Learning based Automatic Speech Recognition (ASR) models are very successful, but hard to interpret. To gain better understanding of how Artificial Neural Networks (ANNs) accomplish their tasks, introspection method…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Clusteringspeech-recognition+1

Deep Neural Networks for Automatic Speaker Recognition Do Not Learn Supra-Segmental Temporal Features

2023-11-01 · Daniel Neururer, Volker Dellwo, Thilo Stadelmann

While deep neural networks have shown impressive results in automatic speaker recognition and related tasks, it is dissatisfactory how little is understood about what exactly is responsible for these results. Part of the…

Speaker Recognition

Discrete Speech Unit Extraction via Independent Component Analysis

2025-01-11 · Tomohiko Nakamura, Kwanghee Choi, Keigo Hojo, Yoshiaki Bando 외

Self-supervised speech models (S3Ms) have become a common tool for the speech processing community, leveraging representations for downstream tasks. Clustering S3M representations yields discrete speech units (DSUs), whi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Clusteringspeech-recognition+1

A CNN-based tool for automatic tongue contour tracking in ultrasound images

2019-07-24 · Jian Zhu, Will Styler, Ian Calloway

For speech research, ultrasound tongue imaging provides a non-invasive means for visualizing tongue position and movement during articulation. Extracting tongue contours from ultrasound images is a basic step in analyzin…

Data Augmentation

Visualizing Deep Neural Networks for Speech Recognition with Learned Topographic Filter Maps

2019-12-06 · Andreas Krug, Sebastian Stober

The uninformative ordering of artificial neurons in Deep Neural Networks complicates visualizing activations in deeper layers. This is one reason why the internal structure of such models is very unintuitive. In neurosci…

speech-recognitionSpeech Recognition