What does a network layer hear? Analyzing hidden representations of end-to-end ASR through speech synthesis
End-to-end speech recognition systems have achieved competitive results compared to traditional systems. However, the complex transformations involved between layers given highly variable acoustic signals are hard to analyze. In this paper, we present our ASR probing model, which synthesizes speech from hidden representations of end-to-end ASR to examine the information maintain after each layer calculation. Listening to the synthesized speech, we observe gradual removal of speaker variability and noise as the layer goes deeper, which aligns with the previous studies on how deep network functions in speech recognition. This paper is the first study analyzing the end-to-end speech recognition model by demonstrating what each layer hears. Speaker verification and speech enhancement measurements on synthesized speech are also conducted to confirm our observation further.
Code (1)
Tasks
Speaker VerificationSpeech Enhancementspeech-recognitionSpeech RecognitionSpeech SynthesisSimilar Papers 제목 키워드 기반
What Information Does a ResNet Compress?
The information bottleneck principle (Shwartz-Ziv & Tishby, 2017) suggests that SGD-based training of deep neural networks results in optimally compressed hidden layers, from an information theoretic perspective. However…
Trainable Reference-Based Evaluation Metric for Identifying Quality of English-Gujarati Machine Translation System
Machine Translation (MT) Evaluation is an integral part of the MT development life cycle. Without analyzing the outputs of MT engines, it is impossible to evaluate the performance of an MT system. Through experiments, it…
Machine TranslationDoes it care what you asked? Understanding Importance of Verbs in Deep Learning QA System
In this paper we present the results of an investigation of the importance of verbs in a deep learning QA system trained on SQuAD dataset. We show that main verbs in questions carry little influence on the decisions made…
Leveraging Deep Representations of Radiology Reports in Survival Analysis for Predicting Heart Failure Patient Mortality
Utilizing clinical texts in survival analysis is difficult because they are largely unstructured. Current automatic extraction models fail to capture textual information comprehensively since their labels are limited in …
Survival AnalysisHeart rate variability code: Does it exist and can we hack it?
Heart rate variability (HRV) has been studied for over 50 years, yet an integrative concept is missing on what HRV's mathematical properties represent physiologically. Here I introduce the notion of HRV code as an attemp…
Heart Rate Variability