The IBM 2015 English Conversational Telephone Speech Recognition System
We describe the latest improvements to the IBM English conversational telephone speech recognition system. Some of the techniques that were found beneficial are: maxout networks with annealed dropout rates; networks with a very large number of outputs trained on 2000 hours of data; joint modeling of partially unfolded recurrent neural networks and convolutional nets by combining the bottleneck and output layers and retraining the resulting model; and lastly, sophisticated language model rescoring with exponential and neural network LMs. These techniques result in an 8.0% word error rate on the Switchboard part of the Hub5-2000 evaluation test set which is 23% relative better than our previous best published result.
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModelingLanguage Modellingspeech-recognitionSpeech RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
The Marchex 2018 English Conversational Telephone Speech Recognition System
In this paper, we describe recent performance improvements to the production Marchex speech recognition system for our spontaneous customer-to-business telephone conversations. In our previous work, we focused on in-doma…
Language ModelingLanguage Modellingspeech-recognitionSpeech RecognitionAdvancing Speech Translation: A Corpus of Mandarin-English Conversational Telephone Speech
This paper introduces a set of English translations for a 123-hour subset of the CallHome Mandarin Chinese data and the HKUST Mandarin Telephone Speech data for the task of speech translation. Paired source-language spee…
TranslationThe IBM 2016 English Conversational Telephone Speech Recognition System
We describe a collection of acoustic and language modeling techniques that lowered the word error rate of our English conversational telephone LVCSR system to a record 6.6% on the Switchboard subset of the Hub5 2000 eval…
Language ModelingLanguage Modellingspeech-recognitionSpeech RecognitionEnglish Conversational Telephone Speech Recognition by Humans and Machines
One of the most difficult speech recognition tasks is accurate recognition of human to human communication. Advances in deep learning over the last few years have produced major speech recognition improvements on the rep…
Language ModelingLanguage ModellingMulti-Task Learningspeech-recognition+1HLT-NUS SUBMISSION FOR 2020 NIST Conversational Telephone Speech SRE
This work provides a brief description of Human Language Technology (HLT) Laboratory, National University of Singapore (NUS) system submission for 2020 NIST conversational telephone speech (CTS) speaker recognition evalu…
Domain AdaptationSpeaker Recognition