paper-with-me

홈 › Papers

Cumulative Adaptation for BLSTM Acoustic Models

2019-06-14 · Markus Kitza, Pavel Golik, Ralf Schlüter, Hermann Ney

This paper addresses the robust speech recognition problem as an adaptation task. Specifically, we investigate the cumulative application of adaptation methods. A bidirectional Long Short-Term Memory (BLSTM) based neural network, capable of learning temporal relationships and translation invariant representations, is used for robust acoustic modelling. Further, i-vectors were used as an input to the neural network to perform instantaneous speaker and environment adaptation, providing 8\% relative improvement in word error rate on the NIST Hub5 2000 evaluation test set. By enhancing the first-pass i-vector based adaptation with a second-pass adaptation using speaker and environment dependent transformations within the network, a further relative improvement of 5\% in word error rate was achieved. We have reevaluated the features used to estimate i-vectors and their normalization to achieve the best performance in a modern large scale automatic speech recognition system.

📄 PDF Abstract BibTeX arXiv:1906.06207

Code (0)

등록된 구현이 없습니다.

Tasks

Acoustic ModellingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Robust Speech Recognitionspeech-recognitionSpeech RecognitionTranslation

Similar Papers 제목 키워드 기반

Deep Recurrent Neural Networks for Acoustic Modelling

2015-04-07 · William Chan, Ian Lane

We present a novel deep Recurrent Neural Network (RNN) model for acoustic modelling in Automatic Speech Recognition (ASR). We term our contribution as a TC-DNN-BLSTM-DNN model, the model combines a Deep Neural Network (D…

Acoustic ModellingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+1

Sequence-level Confidence Classifier for ASR Utterance Accuracy and Application to Acoustic Models

2021-06-30 · Amber Afshan, Kshitiz Kumar, Jian Wu

Scores from traditional confidence classifiers (CCs) in automatic speech recognition (ASR) systems lack universal interpretation and vary with updates to the underlying confidence or acoustic models (AMs). In this work, …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Transformer in action: a comparative study of transformer-based acoustic models for large scale speech recognition applications

2020-10-27 · Yongqiang Wang, Yangyang Shi, Frank Zhang, Chunyang Wu 외

In this paper, we summarize the application of transformer and its streamable variant, Emformer based acoustic model for large scale speech recognition applications. We compare the transformer based acoustic models with …

speech-recognitionSpeech RecognitionVideo Captioning

Progress in Multilingual Speech Recognition for Low Resource Languages Kurmanji Kurdish, Cree and Inuktut

2022-06-01 · LREC 2022 6 · Vishwa Gupta, Gilles Boulianne

This contribution presents our efforts to develop the automatic speech recognition (ASR) systems for three low resource languages: Kurmanji Kurdish, Cree and Inuktut. As a first step, we generate multilingual models from…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Error Reduction Network for DBLSTM-based Voice Conversion

2018-09-26

So far, many of the deep learning approaches for voice conversion produce good quality speech by using a large amount of training data. This paper presents a Deep Bidirectional Long Short-Term Memory (DBLSTM) based voice…

Voice Conversion