paper-with-me

홈 › Papers

First-Pass Large Vocabulary Continuous Speech Recognition using Bi-Directional Recurrent DNNs

2014-08-12 · Awni Y. Hannun, Andrew L. Maas, Daniel Jurafsky, Andrew Y. Ng

We present a method to perform first-pass large vocabulary continuous speech recognition using only a neural network and language model. Deep neural network acoustic models are now commonplace in HMM-based speech recognition systems, but building such systems is a complex, domain-specific task. Recent work demonstrated the feasibility of discarding the HMM sequence modeling framework by directly predicting transcript text from audio. This paper extends this approach in two ways. First, we demonstrate that a straightforward recurrent neural network architecture can achieve a high level of accuracy. Second, we propose and evaluate a modified prefix-search decoding algorithm. This approach to decoding enables first-pass speech recognition with a language model, completely unaided by the cumbersome infrastructure of HMM-based systems. Experiments on the Wall Street Journal corpus demonstrate fairly competitive word error rates, and the importance of bi-directional network recurrence.

📄 PDF Abstract BibTeX arXiv:1408.2873

Code (5)

PaddlePaddle/PaddleSpeech paddle
baidu-research/warp-ctc torch
jb1999/eesen tf
srvk/eesen tf
taozitongxue1/11bee-DeepSpeech tf

Tasks

Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

CNVSRC 2023: The First Chinese Continuous Visual Speech Recognition Challenge

2024-06-14 · Chen Chen, Zehua Liu, Xiaolou Li, Lantian Li 외

The first Chinese Continuous Visual Speech Recognition Challenge aimed to probe the performance of Large Vocabulary Continuous Visual Speech Recognition (LVC-VSR) on two tasks: (1) Single-speaker VSR for a particular spe…

speech-recognitionSpeech RecognitionVisual Speech Recognition

Speech Recognition With No Speech Or With Noisy Speech Beyond English

2019-06-17 · Gautam Krishna, Co Tran, Yan Han, Mason Carnahan 외

In this paper we demonstrate continuous noisy speech recognition using connectionist temporal classification (CTC) model on limited Chinese vocabulary using electroencephalography (EEG) features with no speech signal as …

EEGElectroencephalogram (EEG)General ClassificationNoisy Speech Recognition+2

使用概念資訊於中文大詞彙連續語音辨識之研究 (Exploring Concept Information for Mandarin Large Vocabulary Continuous Speech Recognition) [In Chinese]

2014-12-01 · ROCLINGIJCLCLP 2014 12 · Po-Han Hao, Ssu-Cheng Chen, Berlin Chen
Language Modellingspeech-recognitionSpeech Recognition

EEG based Continuous Speech Recognition using Transformers

2019-12-31 · Gautam Krishna, Co Tran, Mason Carnahan, Ahmed H. Tewfik

In this paper we investigate continuous speech recognition using electroencephalography (EEG) features using recently introduced end-to-end transformer based automatic speech recognition (ASR) model. Our results demonstr…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)EEGElectroencephalogram (EEG)+2

Free Acoustic and Language Models for Large Vocabulary Continuous Speech Recognition in Swedish

2014-05-01 · LREC 2014 5 · Niklas Vanhainen, Giampiero Salvi

This paper presents results for large vocabulary continuous speech recognition (LVCSR) in Swedish. We trained acoustic models on the public domain NST Swedish corpus and made them freely available to the community. The t…

Language Modellingspeech-recognitionSpeech Recognition