paper-with-me

홈 › Papers

Data Augmentation for End-to-end Code-switching Speech Recognition

2020-11-04 · Chenpeng Du, Hao Li, Yizhou Lu, Lan Wang, Yanmin Qian

Training a code-switching end-to-end automatic speech recognition (ASR) model normally requires a large amount of data, while code-switching data is often limited. In this paper, three novel approaches are proposed for code-switching data augmentation. Specifically, they are audio splicing with the existing code-switching data, and TTS with new code-switching texts generated by word translation or word insertion. Our experiments on 200 hours Mandarin-English code-switching dataset show that all the three proposed approaches yield significant improvements on code-switching ASR individually. Moreover, all the proposed approaches can be combined with recent popular SpecAugment, and an addition gain can be obtained. WER is significantly reduced by relative 24.0% compared to the system without any data augmentation, and still relative 13.0% gain compared to the system with only SpecAugment

📄 PDF Abstract BibTeX arXiv:2011.02160

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognitionSpeech RecognitionTranslationWord Translation

Similar Papers 제목 키워드 기반

Improving Code-Switching and Named Entity Recognition in ASR with Speech Editing based Data Augmentation

2023-06-14 · Zheng Liang, Zheshu Song, Ziyang Ma, Chenpeng Du 외

Recently, end-to-end (E2E) automatic speech recognition (ASR) models have made great strides and exhibit excellent performance in general speech recognition. However, there remain several challenging scenarios that E2E m…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationnamed-entity-recognition+7

Improving Code-Switching Speech Recognition with TTS Data Augmentation

2026-01-02 · Yue Heng Yeo, Yuchen Hu, Shreyas Gopal, Yizhou Peng 외 arxiv

Automatic speech recognition (ASR) for conversational code-switching speech remains challenging due to the scarcity of realistic, high-quality labeled speech data. This paper explores multilingual text-to-speech (TTS) mo…

Speech RecognitionData Augmentation

On the End-to-End Solution to Mandarin-English Code-switching Speech Recognition

2018-11-01 · Zhiping Zeng, Yerbolat Khassanov, Van Tung Pham, Hai-Hua Xu 외

Code-switching (CS) refers to a linguistic phenomenon where a speaker uses different languages in an utterance or between alternating utterances. In this work, we study end-to-end (E2E) approaches to the Mandarin-English…

Data AugmentationLanguage IdentificationLanguage ModelingLanguage Modelling+2

Code-Switching without Switching: Language Agnostic End-to-End Speech Translation

2022-10-04 · Christian Huber, Enes Yavuz Ugan, Alexander Waibel

We propose a) a Language Agnostic end-to-end Speech Translation model (LAST), and b) a data augmentation strategy to increase code-switching (CS) performance. With increasing globalization, multiple languages are increas…

Data Augmentationspeech-recognitionSpeech RecognitionTranslation

Language-agnostic Code-Switching in Sequence-To-Sequence Speech Recognition

2022-10-17 · Enes Yavuz Ugan, Christian Huber, Juan Hussain, Alexander Waibel

Code-Switching (CS) is referred to the phenomenon of alternately using words and phrases from different languages. While today's neural end-to-end (E2E) models deliver state-of-the-art performances on the task of automat…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationSequence-To-Sequence Speech Recognition+2