paper-with-me

홈 › Papers

Attention-Based End-to-End Speech Recognition on Voice Search

2017-07-22 · Changhao Shan, Junbo Zhang, Yujun Wang, Lei Xie

Recently, there has been a growing interest in end-to-end speech recognition that directly transcribes speech to text without any predefined alignments. In this paper, we explore the use of attention-based encoder-decoder model for Mandarin speech recognition on a voice search task. Previous attempts have shown that applying attention-based encoder-decoder to Mandarin speech recognition was quite difficult due to the logographic orthography of Mandarin, the large vocabulary and the conditional dependency of the attention model. In this paper, we use character embedding to deal with the large vocabulary. Several tricks are used for effective model training, including L2 regularization, Gaussian weight noise and frame skipping. We compare two attention mechanisms and use attention smoothing to cover long context in the attention model. Taken together, these tricks allow us to finally achieve a character error rate (CER) of 3.58% and a sentence error rate (SER) of 7.43% on the MiTV voice search dataset. While together with a trigram language model, CER and SER reach 2.81% and 5.77%, respectively.

📄 PDF Abstract BibTeX arXiv:1707.07167

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderL2 RegularizationLanguage ModelingLanguage ModellingSentencespeech-recognitionSpeech RecognitionSpeech-to-Text

Similar Papers 제목 키워드 기반

Attention based end to end Speech Recognition for Voice Search in Hindi and English

2021-11-15 · Raviraj Joshi, Venkateshan Kannan

We describe here our work with automatic speech recognition (ASR) in the context of voice search functionality on the Flipkart e-Commerce platform. Starting with the deep learning architecture of Listen-Attend-Spell (LAS…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Voice Quality and Pitch Features in Transformer-Based Speech Recognition

2021-12-21 · Guillermo Cámbara, Jordi Luque, Mireia Farrús

Jitter and shimmer measurements have shown to be carriers of voice quality and prosodic information which enhance the performance of tasks like speaker recognition, diarization or automatic speech recognition (ASR). Howe…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Recognitionspeech-recognition+1

LeVoice ASR Systems for the ISCSLP 2022 Intelligent Cockpit Speech Recognition Challenge

2022-10-14 · Yan Jia, Mi Hong, Jingyu Hou, Kailong Ren 외

This paper describes LeVoice automatic speech recognition systems to track2 of intelligent cockpit speech recognition challenge 2022. Track2 is a speech recognition task without limits on the scope of model size. Our mai…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationDecoder+7

The History of Speech Recognition to the Year 2030

2021-07-30 · Awni Hannun

The decade from 2010 to 2020 saw remarkable improvements in automatic speech recognition. Many people now use speech recognition on a daily basis, for example to perform voice search queries, send text messages, and inte…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

End-to-End Automatic Speech Recognition Integrated With CTC-Based Voice Activity Detection

2020-02-03 · Takenori Yoshimura, Tomoki Hayashi, Kazuya Takeda, Shinji Watanabe

This paper integrates a voice activity detection (VAD) function with end-to-end automatic speech recognition toward an online speech interface and transcribing very long audio recordings. We focus on connectionist tempor…

Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)+2