Free Acoustic and Language Models for Large Vocabulary Continuous Speech Recognition in Swedish
This paper presents results for large vocabulary continuous speech recognition (LVCSR) in Swedish. We trained acoustic models on the public domain NST Swedish corpus and made them freely available to the community. The training procedure corresponds to the reference recogniser (RefRec) developed for the SpeechDat databases during the COST249 action. We describe the modifications we made to the procedure in order to train on the NST database, and the language models we created based on the N-gram data available at the Norwegian Language Council. Our tests include medium vocabulary isolated word recognition and LVCSR. Because no previous results are available for LVCSR in Swedish, we use as baseline the performance of the SpeechDat models on the same tasks. We also compare our best results to the ones obtained in similar conditions on resource rich languages such as American English. We tested the acoustic models with HTK and Julius and plan to make them available in CMU Sphinx format as well in the near future. We believe that the free availability of these resources will boost research in speech and language technology in Swedish, even in research groups that do not have resources to develop ASR systems.
Code (0)
등록된 구현이 없습니다.
Tasks
Language Modellingspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Neural Speech Recognizer: Acoustic-to-Word LSTM Model for Large Vocabulary Speech Recognition
We present results that show it is possible to build a competitive, greatly simplified, large vocabulary continuous speech recognition system with whole words as acoustic units. We model the output vocabulary of about 10…
Language ModelingLanguage Modellingspeech-recognitionSpeech RecognitionEnd-to-End Attention-based Large Vocabulary Speech Recognition
Many of the current state-of-the-art Large Vocabulary Continuous Speech Recognition Systems (LVCSR) are hybrids of neural networks and Hidden Markov Models (HMMs). Most of these systems contain separate components that d…
Acoustic ModellingLanguage ModelingLanguage Modellingspeech-recognition+1Comparison of Lattice-Free and Lattice-Based Sequence Discriminative Training Criteria for LVCSR
Sequence discriminative training criteria have long been a standard tool in automatic speech recognition for improving the performance of acoustic models over their maximum likelihood / cross entropy trained counterparts…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)GPULanguage Modeling+3First-Pass Large Vocabulary Continuous Speech Recognition using Bi-Directional Recurrent DNNs
We present a method to perform first-pass large vocabulary continuous speech recognition using only a neural network and language model. Deep neural network acoustic models are now commonplace in HMM-based speech recogni…
Language ModelingLanguage Modellingspeech-recognitionSpeech RecognitionThe RWTH Aachen LVCSR system for IWSLT-2016 German Skype conversation recognition task
In this paper the RWTH large vocabulary continuous speech recognition (LVCSR) systems developed for the IWSLT-2016 evaluation campaign are described. This evaluation campaign focuses on transcribing spontaneous speech fr…
Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition