Data Augmentation for End-to-end Code-switching Speech Recognition
Training a code-switching end-to-end automatic speech recognition (ASR) model normally requires a large amount of data, while code-switching data is often limited. In this paper, three novel approaches are proposed for code-switching data augmentation. Specifically, they are audio splicing with the existing code-switching data, and TTS with new code-switching texts generated by word translation or word insertion. Our experiments on 200 hours Mandarin-English code-switching dataset show that all the three proposed approaches yield significant improvements on code-switching ASR individually. Moreover, all the proposed approaches can be combined with recent popular SpecAugment, and an addition gain can be obtained. WER is significantly reduced by relative 24.0% compared to the system without any data augmentation, and still relative 13.0% gain compared to the system with only SpecAugment
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognitionSpeech RecognitionTranslationWord TranslationSimilar Papers 제목 키워드 기반
Improving Code-Switching and Named Entity Recognition in ASR with Speech Editing based Data Augmentation
Recently, end-to-end (E2E) automatic speech recognition (ASR) models have made great strides and exhibit excellent performance in general speech recognition. However, there remain several challenging scenarios that E2E m…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationnamed-entity-recognition+7Improving Code-Switching Speech Recognition with TTS Data Augmentation
Automatic speech recognition (ASR) for conversational code-switching speech remains challenging due to the scarcity of realistic, high-quality labeled speech data. This paper explores multilingual text-to-speech (TTS) mo…
Speech RecognitionData AugmentationOn the End-to-End Solution to Mandarin-English Code-switching Speech Recognition
Code-switching (CS) refers to a linguistic phenomenon where a speaker uses different languages in an utterance or between alternating utterances. In this work, we study end-to-end (E2E) approaches to the Mandarin-English…
Data AugmentationLanguage IdentificationLanguage ModelingLanguage Modelling+2Code-Switching without Switching: Language Agnostic End-to-End Speech Translation
We propose a) a Language Agnostic end-to-end Speech Translation model (LAST), and b) a data augmentation strategy to increase code-switching (CS) performance. With increasing globalization, multiple languages are increas…
Data Augmentationspeech-recognitionSpeech RecognitionTranslationLanguage-agnostic Code-Switching in Sequence-To-Sequence Speech Recognition
Code-Switching (CS) is referred to the phenomenon of alternately using words and phrases from different languages. While today's neural end-to-end (E2E) models deliver state-of-the-art performances on the task of automat…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationSequence-To-Sequence Speech Recognition+2